Handwritten Text Recognition of Historical Argentinian Census Documents

Hämtar...
Bild (thumbnail)

Publicerad

Författare

Typ

Examensarbete för masterexamen
Master's Thesis

Modellbyggare

Tidskriftstitel

ISSN

Volymtitel

Utgivare

Sammanfattning

There is a lack of digitized historical census records, especially in the Spanish language. In this project the large language model (LLM) Chandra was used for handwritten text recognition on the 1895 Argentinian census. The census lacks ground truth transcriptions so an iterative data annotation process was used to create a training dataset on which the LLM was fine tuned. To further reduce the error rate image transformations were used. Testing concluded that a median filter is able to reduce the error rate by about 2,5 percentage points. Another LLM from the Qwen3 series was trained to correct the output from the transcription model, but due to a lack of training data this method did not perform well. The final model was evaluated on a testing set and scored an average CER of 12%, WER of 8% and FER of 7%. The result is decent but future work on the model should focus on reducing the error rate of the name column in the dataset. The project has resulted in 102 transcribed pages from the 1895 Argentinian census and a fine tuned LLM which can be used to transcribe further pages in the dataset. The model can also be used as a foundation for other similar digitization projects. Keywords:

Beskrivning

Ämne/nyckelord

HTR, LLM, OCR, HDP, census, Spanish, Argentinian

Citation

Arkitekt (konstruktör)

Geografisk plats

Byggnad (typ)

Byggår

Modelltyp

Skala

Teknik / material

Index

Endorsement

Review

Supplemented By

Referenced By