Large Language Models are not a One-Size-Fit All Solution for Archives and Records Management
DOI:
https://doi.org/10.5555/axjn5606Keywords:
Archival Description, Large Language Models, Archives and Records Management, Algorithmic Bias, Professional Ethics, Generative AIAbstract
The proliferation of Large Language Models (LLMs) in the public sector promises significant gains in efficiency and improved service delivery. However, the use of LLMs, especially in low-resource settings, is often unguided and informal. The informal adoption of off-the-shelf Large Language Models (LLMs) by archival practitioners using personal accounts, without institutional oversight, fine-tuning, or structured prompt engineering, poses risks to professional practice. While these tools offer a practical means of addressing backlogs in archival processing, novice adopters may not be aware of the risks associated with using decontextualised models. This study employed a novice-driven experimental case study to evaluate three commercially available LLMs: DeepSeek, ChatGPT, and Gemini, across five archival records from the Botswana National Archives. Using a minimalist prompt with no domain-specific instructions, the outputs were assessed against eleven descriptive categories derived from archival principles. The findings reveal systematic risks: hallucination of non-existent information that undermines provenance integrity; unsubstantiated temporal assertions that misrepresent chronological relationships; non-compliance with descriptive standards, thus, preventing interoperability; variable vision capabilities introducing collection-specific biases; and inconsistent handling of sensitive historical language. No single LLM proved universally superior, indicating that domain-specific fine-tuning must consider multimodal and multi-agent capabilities. The originality of this study lies in translating empirical findings into actionable professional guidance, including a Three-Level LLM Adoption Framework (Assisted Exploration, Conditional Automation, Integrated Assistance), an implementation matrix for low-resourced settings, and a decision flowchart. The study concludes that safe LLM integration requires procedural rigour, shared resources, and human verification of every output before it enters the archival record, offering a realistic pathway between uncritical automation and technological abstinence.
Discussion
No comments yet.