A budget Nikon hanging over a table, cheap LED bulbs, a glass sheet pulled from a photocopier. That rig belonged to Ibteda Digital Library, a small community archive in Pakistan. The books it set out to save, the members bought with their own money.12
For nearly a decade, pressing the shutter was the easy part. Somebody still had to open every shot in Photoshop and squeeze two finished pages out of it: crop, straighten, choose the margin, decide what to do with a torn corner. One click cost nothing. Do it across hundreds of pages and you are spending years.1
In April 2026 the team suspended daily work. 526,392 dual-page photographs were still unfinished. By hand they would have taken years more.12
Then the project turns odd. The thing they reused to take over was not bought software. It was their own retouching — 575,729 hand-finished pages, now acting as the labels of a computer-vision pipeline.1
One click at a time
The archive was born of frustration. Around 2015, putting together an edition of Ghalib's Diwan for the 150th anniversary of his death, one founder saw how unreachable the historic editions were: out of print, locked in university collections, parked on private shelves. Ibteda was the answer — find the physical books and build the archive yourself.1
Production ran in three stages. Capture: photograph an open spread through the glass. Post-process: an operator finishes both pages in Photoshop. Publish: store two ordered, consistently sized page images.1 The catalogue now passes 5,000 books; the analyzed inventory holds 822,161 dual-page photographs, roughly 1.64 million page sides.1
The material is brutal. Nineteenth-century lithographs sit beside mid-century prose, dictionaries, magazines, pages covered in marginal notes, illustrated volumes. A rule tuned for neat Urdu prose would fail on a large part of the shelf.12
Urdu handwriting, often in Nastaliq script, complicates digitization in a concrete way: a lot of dots, diacritics and tiny symbols. Dust, dirt and blemishes blend easily with real ink. Every book becomes its own corner case.2

Finished pages became labels
For a long time, capture outpaced post-processing: new photos piled up faster than they could be finished. One member first tried classical computer vision — contour detection, projection profiles, paper-color thresholds. Each approach worked on a narrow slice of books and failed on the rest. Gutter shadows and curved bindings produced convincing false edges. OpenCV stayed useful, but no fixed rule could pick the project's preferred crop.12
The turn came from looking differently at the 575,729 Photoshop files already produced. They stopped being publication files and started being a decade of cropping decisions: rotation, margins, physical side, and a book-specific framing repeated across hundreds of pages.1
Each finished page still had to be linked back to its raw spread, with the geometry the operator chose reconstructed. The author ran SIFT against nearby candidates, filtered descriptors, then fitted a homography. Crucially, the hard part was rejection, not matching: dense Urdu print repeats strokes, words, borders and columns, so a neighboring page can produce a transformation that looks convincing at a glance. The team set conservative thresholds — a false label is worse than losing a training page.1
Of the 36,450 correspondences attempted, 35,648 were accepted: 97.8%. Interior pages reached 98.6%; near the book's edges, where covers and one-sided leaves live, only 77.9%.1
Ten crops beat a bigger model
The first geometry model stayed deliberately simple: a 704×704 input, an ImageNet-pretrained ResNet-34 trunk, and two heads for the physical-left and physical-right pages. Each side's output collapses to six values — center, width, height, rotation — then decodes to four corners. The loss is a Smooth L1 across the corners, in the final page's pixels.1
The decisive insight hides in what the camera does not show. The photos contain the physical paper edge, but the archive's chosen inset is invisible. One book looks right with a 40-pixel inset, another with 120. Trained on old books, the network cannot guess the preference of a volume it has never seen.1
So the author stopped asking the network to guess. The Studio picks five spreads through the book; an operator corrects both pages, giving ten correction vectors. The median of those ten residuals is stored as a correction for the volume. Across 249 books in a confirmation cohort, the ten-crop correction lifted the pass@80 rate from 0.7107 to 0.8271.1
The counterintuitive part is the cleanest: adding data or growing the model did not improve quality on unseen books. Trained on more books, the larger ResNet-50 actually came out slightly worse after calibration (0.8138 vs 0.8271). The inset belongs to each book, so more volumes blurred the pattern instead of tightening it.12
Do not erase the diacritics
Geometry solved, the uglier question stayed: which marks to remove without touching the printed page. The naive approach — subtract the finished page from the raw one — was unusable. Tiny alignment errors lit up both sides of every Urdu stroke, and ordinary contrast changes looked like removals. On one measured page, the naive mask covered 7.28% of the image, 49.2% of which was ink the team had kept.1
Something finer than "remove or keep" was needed. The team built three label states: REMOVE for a region the operator removed completely, KEEP for legitimate ink, and IGNORE for ambiguous pixels excluded from the loss. Urdu dots and diacritics became explicit hard negatives: small dark marks the network must absolutely preserve.1
The retouch network stays binary — one removal logit per pixel — but the label decision is careful. A stricter label cut reached 0.6024 IoU on the mark tier with zero measured false-positive diacritic pixels; an inclusive cut gained a little IoU but produced 161 false-positive diacritic pixels per patch. The version that protects the language won.1
The network never acts on its own. It proposes a support mask; component rules then decide whether the system acts or abstains. Below a component size and confidence threshold, the mark is left alone — half a stamp or half a dark band often looks worse than leaving it. Paper is reconstructed with ordinary OpenCV and NumPy, and the raw page is copied back exactly outside the declared support.1
Stopping before proving everything
The team also documents what did not work. Raising the training books from 378 to 572, lifting the input to 1024 pixels, a spatial head, higher resolution: none improved quality on unseen books. The generalization wall is owned, not hidden.1
The write-up even owns its mistakes. For a long stretch, one random calibration draw seemed to show a significant difference between two models, when across 20 draws the effect vanished. That candor makes the rest easier to trust: the published thresholds are Ibteda's own operating points, not a portable guarantee for other archives.1
At the April 2026 stop, the pipeline still needed ten calibration crops per new book and roughly a third of page sides went to review at high-quality operating points. The output is not free; it has simply become measurable, stoppable, inspectable and resumable later.1

More than 1,300 books already online
The important part is what is already published. More than 1,300 finished books are accessible on the Internet Archive, including over 350 lithographed volumes, nearly 500 poetry collections and more than 1,000 periodical issues, with holdings dating back to 1821.13
The raw photographs survive on a ZFS storage pool protected by BLAKE3 manifests and integrity audits, together with 35,648 accepted geometry labels and the finishing Studio itself.1 The heart of the story sits in the closing line: "we turned a decade of human effort into a process that can be measured, stopped, inspected, and resumed whenever GPU time and operator attention are available."1
The team notes its own final irony: they began frustrated with private collectors, and ten years later they had bought over 5,000 volumes themselves and built new bookshelves. "Only this time, the bits are safe."1
