Thomson Reuters has officially launched Thomson, the company’s first proprietary large language model, developed in-house.
Frontier labs have typically spent billions of dollars on compute and years of infrastructure investment to reach the frontier. Thomson Reuters started from an open-source foundation and invested $40 million to train Thomson into the right intelligence for the jobs that matter most, covering talent and compute. The result is a model Thomson Reuters fully controls, without the heavy inference costs of typical frontier models.
Thomson also uses mid-training and post-training techniques, drawing on content from Westlaw, Practical Law, Checkpoint, and Reuters, with hundreds of subject matter experts integrated from the design of training objectives through to the final evaluations. The model has been trained on less than 10% of Thomson Reuters content so far, and what comes next is not simply feeding it more data. It is continued discovery of new kinds of specialization and understanding, made possible only by building on decades of proprietary content and editorial expertise.
Evaluations of Thomson’s underlying foundation model are available in the technical report about the model’s development. Ahead of this launch, Thomson Reuters began opening the model to a group of legal and AI academics for direct evaluation.
Thomson Reuters will continue to make the model available to external parties to aid in the further validation and development of Thomson over the coming weeks and months. Thomson Reuters is also making a “small” version of Thomson available as an open-weight model on Hugging Face for academic and non-commercial use to further aid in this validation.