---
title: "Lab Update: Reading Faded Ink"
description: "How we fine-tuned vision-language models and Kraken CTC to transcribe 18th-century Barbadian secretary hand records."
pubDate: 2026-07-17
author: "Anaiya"
audio: "/audio/2026-07-17-zindi-breakthrough.mp3"

---

> *"All that you touch You Change. All that you Change Changes you. The only lasting truth is Change."*
>
> *Octavia E. Butler*

![Change and Transformation Abstract](/images/blog/zindi.jpg)

Wuh gine on! Anaiya here. 👋🏾

It has been an intense week in the lab. ⏱️ We jumped headfirst into the **R.O.A.D. Barbados Historic Handwriting Challenge** on Zindi, an initiative aimed at transcribing thousands of degraded 18th- and 19th-century Barbadian archival records.

When you look at these colonial-era survey ledgers and legal deeds, you quickly realize how tough the challenge is. You are dealing with historic **secretary hand**: multiple scribes with wildly different cursive flourishes, faded iron gall ink, water stains, bleed-through from the back of the page, and severe line overlap. Off-the-shelf optical character recognition baselines struggle significantly on documents like these, often producing character error rates exceeding 60%.

When we ran our initial zero-shot vision baselines, our scores were stuck in the low 0.30s. The models were hallucinating modern English phrasing instead of deciphering archaic abbreviations and orthography.

We locked in. 🔬

1. **Fixing the Metric Trap:** We caught an inverted scoring trap early where we were calculating error rate inversely. Zindi's official leaderboard score is normalized as `1 − (0.5·wCER + 0.5·wWER)` where **higher is better**. Aligning our local 410-line holdout to match the exact competition metric brought our local validation within 0.005 of our subsequent leaderboard submission.
2. **Vision-Tower LoRA:** We discovered that standard PEFT starter configs only targeted language attention projections (`q_proj`), completely bypassing Qwen's visual modules (`proj`, `fc1`, `fc2`). Enabling vision-tower adaptation contributed to a +0.064 jump on the public leaderboard (as recorded in our July 2026 submission logs).
3. **Dual-Engine Strategy:** We deployed a two-pronged architecture: fine-tuning **Qwen2-VL-2B** on the language and vision side, while training **Kraken CTC** (fine-tuned from the McCATMuS historical base) inside a WSL2 Ubuntu environment.

That combination propelled us from rank 121 up to **rank 59** with a public leaderboard score of **0.8717** (July 16, 2026 submission snapshot).

This sprint also forced us to automate our workspace discipline: strict holdout isolation, deterministic preprocessing transforms, and reusable challenge scaffolding scripts.

We are building the tools that build the tools. 🛠️

Talk soon,  
**Anaiya ✨**
