portfolio.cases

    Case studies

    Three chapters of the same system: rebuilding localization as AI infrastructure, proving it on a 45,000-word website launch, and removing the last manual steps from the pipeline.

    case_01.ai_pipelinelive · jan 2025

    Rebuilding localization as an AI-driven pipeline

    The flagship: replacing an agency-dependent, 4-day translation cycle with an AI pipeline that delivers in minutes, without letting quality become an act of faith.

    −80%
    cumulative cost vs. agency baseline¹
    10 min
    typical delivery, down from 4 days
    18
    languages, all AI-translated since Jan 2025
    1 tool
    replaced 5 years of parallel spreadsheet workflows

    Before — 2023

    • Weekly agency cycles; ~4-day turnaround per request
    • Two content designers as the mandatory middle layer for every team
    • Jira tickets, handovers and "who owns this key?" confusion
    • QA only when adding a new language — expensive, inconsistent
    • Translation spend scaling linearly with volume

    After — 2025

    • AI translation for all 18 in-product languages, minutes to delivery
    • Key creation owned by product trios via the Figma ↔ Lokalise plugin
    • Staging/production branching; automated daily staging deliveries
    • Continuous AI LQA + quality scoring on every batch
    • ~80% lower cumulative cost while volume grew 30%+

    The decision wasn't "AI: yes/no" — it was risk allocation

    I modelled three rollout proposals against cost, risk and market coverage:
    • AI for everything (~80% savings, highest risk),
    • AI with linguists kept for core markets (~50%),
    • and a conservative hybrid weighted by users per language (~33–35%).
    I recommended starting conservative and earning our way toward full AI with quality data — which is exactly what happened over the following 15 months.

    Quality was benchmarked, not assumed

    Same content, translated in parallel by AI and by professional linguists, scored with the same LQA rubric. The result made the business case: AI error rates were statistically comparable to human linguists — and the error types (terminology, glossary adherence) were exactly the ones fixable through better glossaries, style guides and source text.

    AI translations — major errors14 / 43 (32.5%)
    Professional linguists — major errors12 / 40 (30%)

    Fix causes, not symptoms

    Iterating on the LQA loop taught us that editing bad translations one by one is treating the symptom. The durable fixes were upstream: glossary clean-up, style guides per language, context on keys, and better source text. Review time per batch dropped from ~45 to ~30 minutes as the loop matured.

    Built to fail safely

    Every stage of the rollout had a reversal path: branch-gated releases, tagged AI batches, translation-memory rollback. "A safe way to fail" wasn't a slogan — it's what made it possible to move fast on 18 live languages without betting the product on an unproven model.

    ¹ metric provenance:
    * cost figures compare monthly translation spend against the pre-AI agency-only budget baseline (2023);
    * 57% was the directly measured savings at the Sept 2025 audit;
    * ~80% is the cumulative reduction as remaining agency review was phased out. Turnaround range (60–99%) depends on task type;
    * "10 minutes" reflects standard AI-translated product keys. Quality benchmark: parallel AI vs. linguist translation of identical content, scored with the same LQA rubric, 2024.