Large language models (LLM) have quickly become a standard tool for extracting quantitative metrics from unstructured text. Are these metrics more like new data or a biased filter? In the July 2026 revision of their paper entitled "LLMs and Systematic Measurement Error", William Grieser, Anthony Cookson, Maryam Fathollahi and Eshwar Venugopal examine how LLMs process text and thereby introduce errors of omission and injection. The authors then examine how LLM omission and injection rates influence summaries of firm risk disclosures and earnings call transcripts. Using four widely used models (GPT-4o-mini, Grok-3-mini, Phi-4, Llama-3.3-70B) to summarize annual Form 10-K, Item 1A risk factor disclosures and quarterly earnings call transcripts of S&P 500 firms during 2005 through 2024, they find that:
Subscribe to Keep Reading
Get the research edge serious investors rely on.
- 1,200+ research articles
- Monthly strategy signals
- 20+ years of backtested analysis
Cancel anytime