About Wurdahuzdą

Proto-Germanic Dictionary and Etymological Database

wurdahuzdą ᚹᚢᚱᛞᚨᚺᚢᛉᛞᚨ /ˈwur.ðɑ̃.ˌxuz.dɑ̃/ n. ‘word-hoard; treasury or store of words’

The logo plays on the two runes corresponding to the first elements of the compound: Wunjo (ᚹ) for wurdą (‘word’) and Hagalaz (ᚺ) for huzdą (‘hoard, treasure’). The Wunjo is mirrored and inverted so that the two forms combine into a stylized Hagalaz.

The earliest securely datable occurrence of the compound in a Germanic language is Old English wordhord:

Ða se Wisdom eft wordhord onleac …
(Then Wisdom again unlocked [the] word-hoard …)

It occurs in the mid-tenth-century manuscript of the Metres of Boethius (Metre 6.1), and the compound is also found in Widsith, Beowulf, Andreas, Vainglory, and The Order of the World, works whose dates of composition do not permit us to establish a securely earlier occurrence.

The compound does not appear to have survived elsewhere in the Germanic family and may simply have been an Old English creation. Nevertheless, word-hoard seemed an irresistible name for this project: a treasury of words reconstructed for the common Germanic language from which English, German, Dutch, the Scandinavian languages, Gothic, and the other Germanic languages ultimately descend. I hope historical linguists will forgive this one deliberately anachronistic Proto-Germanic reconstruction.

The story of how the dictionary came to be is considerably longer.

My interest in Germanic languages began when I was a child learning Pennsylvania Dutch. I was fascinated by the discovery that there was, in effect, another way to speak to my great-grandmother, whose native language it was. German came easily to me, and in high school in the 1980s I spent a formative year living with a German family and studying at a Gymnasium in Mönchengladbach, in what was at the time West Germany (I also had the opportunity to learn a Low German dialect and make frequent visits to the Netherlands). I later spent another year at the Universität zu Köln while double-majoring in German and Finance at Pennsylvania State University. At the time I imagined a career in international finance. Graduating shortly after the market crash of the late 1980s made that plan considerably less attractive.

I stayed in school, completed a master's degree in Information Systems, and spent roughly a decade working in information technology and consulting in Washington, D.C. After losing an acquaintance in the attacks of September 11, 2001, I reconsidered what I wanted to do with my life and returned to graduate school to study the subject that had truly been my life-long passion: Germanic linguistics.

I was fortunate to receive a six-year fellowship in the Germanic Linguistics MA/PhD program at the University of California, Berkeley, where I studied comparative and historical linguistics under Irmengard Rauch and Thomas Shannon. For me, it was the realization of a long-held ambition. A family emergency eventually required me to leave California before completing the PhD, but the training profoundly changed the course of my professional life.

In the years that followed, I published several books on language and linguistics, which together sold more than 20,000 copies, including a history of the English language that became the most widely read of them. The combination of linguistic training and an earlier background in computing also led me into computational linguistics. I eventually worked full-time and on consulting projects for organizations including Oracle, Grammarly, and Apple. By the late 2010s, computational linguistics had gone from an unusually difficult profession to explain to people to a rapidly expanding field. I retired in late 2024 after a three-year project with Apple, one of the most rewarding experiences of my working life.

Retirement left me with an obvious question: what next? I knew fairly quickly that I wanted another substantial writing and research project, and etymology had always been one of my favorite areas of linguistics.

Wurdahuzdą began with the Proto-Germanic material in Wiktionary. I initially imagined that I could extract the data, regularize it, expand it somewhat, and turn it into a more systematic reference work. I soon discovered that the task was much larger. Wiktionary contains an extraordinary amount of material, assembled through many years of volunteer work, and this project could not exist without it. At the same time, its Proto-Germanic entries vary considerably in depth, documentation, reconstruction, and bibliographic support. Many etymologies depend primarily on a single older reference, while others reflect analyses assembled incrementally by different editors over many years.

That is entirely understandable for a collaborative encyclopedia. It also suggested an opportunity for a different kind of resource: a work in which the Proto-Germanic lexicon could be reviewed systematically, entry by entry, against the specialist literature; competing analyses could be made explicit; newer scholarship could be incorporated; reconstructions and notation could be standardized; and the evidence behind editorial decisions could be documented consistently.

What began as an editing project soon became a new dictionary.

The scale is substantial: more than 5,300 reconstructed headwords, along with their meanings, etymologies, morphology, descendants, cognates, and supporting bibliography. Each etymology is being individually researched and revised, with particular attention to Proto-Indo-European reconstruction and to developments in Germanic historical phonology and morphology. The aim is not to replace the scholarship on which the dictionary depends, but to bring that scholarship together in a form that is current, transparent, searchable, and useful for further research.

I also wanted the project to be more than a book. Because the dictionary is database-driven, the same underlying data can support tools that a conventional printed dictionary cannot easily provide: structured searches across grammatical and etymological categories, inflected-form analysis, morphological searches, reverse lookup, comparisons across descendant languages, downloadable datasets, and statistical views of the lexicon. The web application and the dictionary are therefore two presentations of the same scholarly resource rather than separate projects.

This combination reflects the two halves of my own career. Historical linguistics taught me to think about how words and languages change across centuries; computational linguistics taught me to think about what becomes possible when linguistic information is represented as structured data. Wurdahuzdą is where those two paths finally meet.

I hope the result will be useful at several levels: to specialists in Germanic and Indo-European linguistics; to students learning how reconstruction and etymological argument work; to lexicographers and digital-humanities researchers who want structured historical-language data; and to anyone who is simply curious about where the Germanic vocabulary came from.

The project is intended to remain openly available so that the work invested in it can continue to circulate, be tested, corrected, reused, and built upon. Wiktionary's editors freely contributed the material that made this project possible; one of my goals is to return something substantially expanded to that same larger community of scholars, students, developers, and language enthusiasts.

I have the unusual luxury, in retirement, of being able to devote the time to a project of this scale. I expect it to occupy me for years. That does not particularly trouble me. I have come to think of Wurdahuzdą as the project that brings together nearly everything I have done professionally, and, perhaps, as the work I most hope to leave behind.

If you are a scholar, editor, publisher, digital-humanities researcher, or developer interested in the project—or if you have an idea for a search, visualization, dataset, or other tool that these data could support—I would be very glad to hear from you: sshay /at/ berkeley d0t edu.