Handwritten Text Recognition of Historical Argentinian Census Documents
Hämtar...
Ladda ner
Publicerad
Författare
Typ
Examensarbete för masterexamen
Master's Thesis
Master's Thesis
Modellbyggare
Tidskriftstitel
ISSN
Volymtitel
Utgivare
Sammanfattning
There is a lack of digitized historical census records, especially in the Spanish language.
In this project the large language model (LLM) Chandra was used for handwritten
text recognition on the 1895 Argentinian census. The census lacks ground
truth transcriptions so an iterative data annotation process was used to create a
training dataset on which the LLM was fine tuned. To further reduce the error rate
image transformations were used. Testing concluded that a median filter is able
to reduce the error rate by about 2,5 percentage points. Another LLM from the
Qwen3 series was trained to correct the output from the transcription model, but
due to a lack of training data this method did not perform well. The final model
was evaluated on a testing set and scored an average CER of 12%, WER of 8% and
FER of 7%. The result is decent but future work on the model should focus on
reducing the error rate of the name column in the dataset. The project has resulted
in 102 transcribed pages from the 1895 Argentinian census and a fine tuned LLM
which can be used to transcribe further pages in the dataset. The model can also
be used as a foundation for other similar digitization projects.
Keywords:
Beskrivning
Ämne/nyckelord
HTR, LLM, OCR, HDP, census, Spanish, Argentinian
