Skip to content
WM William Mattingly
Home Talks Projects About CV Blog Contact
Menu
Home Talks Projects About CV Blog Contact

Writing

Blog

Notes, essays, and tutorials on machine learning, NLP, and the digital humanities.

  • October 6, 2026

    Who Is “Boom BK”? Linking 7.2 Million Catalogued Names to Wikidata with a Small Model

    • entity linking
    • wikidata
    • europeana
    • small models
    • gliner
    • digital humanities

    How a 340-million-parameter model, trained on data labelled with the help of a larger one, decided which Wikidata person each of 7.2 million names in Europeana's records refers to, or that none of them does. Built step by step, with the decisions about caution, data, and scale explained along the way, and run end to end on Yale's cluster without a single API call.

    Read post →
  • September 24, 2026

    Where the Readings Disagree: Finding HTR Errors by Asking a Model Five Times

    • htr
    • manuscripts
    • vision-language models
    • qwen
    • digital humanities

    A small local app that runs a handwritten text recognition model several times over the same manuscript page and uses the disagreement between runs to show, word by word and letter by letter, where the errors are likely to be. Tested on Caroline minuscule, and on Old English, a language the model never saw.

    Read post →
  • June 25, 2026

    Reading the Cluster at a Glance with Yale SLURM Utils

    • yale
    • hpc
    • slurm
    • python

    A small command-line tool that turns dense SLURM output into a readable, live dashboard — and why that matters when you share a supercomputer to train and run large language models.

    Read post →
  • May 18, 2026

    Dissolving a 398-Node Geographic Cycle with Gemini Flash Lite

    • yale
    • graph theory
    • lux

    How we found, visualized, and automatically fixed a massive circular reference in the LUX places hierarchy — for about three cents.

    Read post →
  • May 11, 2026

    Whisper Achieves 85% Accuracy on Holocaust Testimonies in Yale's Fortunoff Archive

    • asr
    • machine learning
    • holocaust

    We evaluated Whisper on 1,847 Holocaust testimonies and found 85 percent accuracy, though the model routinely normalizes raw speech and heritage spellings.

    Read post →
  • May 3, 2026

    Parsing 3.6 Million Historical Names with Small Models

    • yale
    • machine learning
    • name parsing

    We moved from expensive frontier AI to fine-tuned Qwen 3.5 models to parse historical data, achieving 96% accuracy by switching from JSON to YAML.

    Read post →
  • April 25, 2026

    Bye Conda, Hello uv

    • uv
    • environment management
    Read post →

William Mattingly, PhD · 2026

  • GitHub
  • X
  • LinkedIn
  • YouTube
  • Hugging Face