In short
Thomson Reuters claims that its own model, Thomson, competes with Claude Opus 4.8 and outperforms GPT-5.5—at a fraction of their size and cost. However, all benchmarks are internal, and the main argument is based on proprietary data that competitors do not have.
The market had assumed that all it would take was a powerful general-purpose model with RAG integrated into it. Thomson Reuters took a different approach: building its own LLM tailored for legal work and training it on proprietary content unavailable to competitors. According to the company, Thomson’s model competes with Claude Opus 4.8 and outperforms GPT-5.5, Claude Sonnet 5, and Gemini 3.1 Pro.
The main point: general-purpose models are good for everyone, but professionals need accuracy, not versatility. For a lawyer, “almost correct” is a mistake with consequences. In 2024, Thomson Reuters acquired Safe Sign Technologies and began building its own model instead of relying on others’.
Thomson is built on an open-source foundation and fine-tuned using content from Westlaw, Practical Law, Checkpoint, and Reuters. Hundreds of experts evaluated the model’s conclusions, identified errors, and verified that it reasons like a lawyer, not like a general-purpose assistant. Client data is not used for training.
The key difference from frontier models is access to proprietary data. In a test involving 53 legal queries written by in-house experts, Thomson—with access to Westlaw and Practical Law—demonstrated greater completeness and factual accuracy than frontier models with access to the web. The model cites sources that can be verified. This isn’t architectural magic—it’s data that OpenAI and Anthropic don’t have.
But critically: all benchmarks are internal and self-reported. The methodology is “LLM as judge,” calibrated against assessments by Thomson Reuters experts. The model has not yet been launched. The first application—Tabular Analysis at CoCounsel Legal in August—involves structured document analysis with a clear accuracy standard. This is the safest way to start: a large-scale, predictable task where the domain model’s advantage is easy to measure and difficult to dispute.
Less than 10% of Thomson Reuters’ content was used. If the model already claims parity with the state-of-the-art using just one-tenth of the data, the question is what will happen with the remaining 90%. Either the advantage is greatly exaggerated, or domain specialization is truly a game-changer.
The strategic takeaway is simpler than the benchmarks: if you have proprietary data and domain expertise, building your own model may be more profitable than relying on someone else’s. Thomson Reuters isn’t claiming that general-purpose models are useless—it’s arguing that, when high stakes are at play, specialization trumps versatility. This challenges the narrative that “one big model will solve everything.”
Until independent tests and real-world applications are available, this is a promise, not a fact. But the direction—domain-specific models trained on proprietary data—seems like a more sustainable trend than yet another race for scale.