Most companies access AI models such as ChatGPT and Claude through the cloud. Due to security concerns, many companies therefore route requests through an AI gateway, which monitors usage and removes any personally identifiable information from prompts before they are passed through to the public cloud. This is easy to implement but can easily result in unpredictable token costs, as many companies have already discovered. Uber spent its annual AI coding token budget in just four months, while a report from Asana found that 82% of companies had experienced unexpected AI cost increases. One unnamed company reportedly ended up with a bill of $500 million.
Concerns over costs and security have led some enterprises to run AI models in a private cloud, which involves dedicated servers for a single customer. This may involve renting hardware from companies like CoreWeave or Lambda Labs, or using dedicated isolated services from public cloud vendors, such as Amazon’s AWS GovCloud, where dedicated server resources are set up to run in a specific region. The third way is to run AI models locally in a corporate data centre, using dedicated GPU clusters managed by data centre staff. According to Broadcom, 48% of enterprises run their AI on private clouds, 41% on public clouds, 8% on their own data centres, with 3% running elsewhere, such as in edge environments like a factory floor.
All of these assume that companies run commercially available AI models, usually supplemented by corporate data via the retrieval-augmented generation (RAG) technique.
Some enterprises have gone further and decided to train their own dedicated large language models. A recent example is Thomson Reuters, which in August 2026 announced that it had completed a proprietary, in-house large language model, based on an open-weight model. This model has been trained on a bit less than 10% of Thomson Reuters content, including Westlaw and Practical Law. The model is now being tested on various legal questions, with good early feedback compared to frontier models like Claude, at least according to Thomson Reuters. Training your own LLM is a major commitment: Thomson Reuters invested $40 million in staff and compute to develop the model.
The model actually runs on a private cloud operated by Lambda Labs and another cloud AI company called Together AI. It is based on the open-weight Chinese model Qwen from Alibaba. This is a relatively small 35-billion parameter mixture-of-experts model, with around 3 billion parameters active for each token. By comparison, Claude Sonnet is estimated to have about 1 trillion parameters. Thomson Reuters is not the first enterprise to produce a proprietary LLM. Bloomberg built a 50-billion parameter financial LLM trained on its proprietary financial data back in 2023. Samsung also built its own model, called Gauss. JPMorgan took a different route, with its “LLM Suite” actually based on proprietary data and multiple existing LLMs. The Thomson Reuters model shows that open-weight LLMs have now advanced sufficiently that a company can consider building its own proprietary LLM based on them, albeit at a substantial training cost. This is not going to be appropriate or worthwhile for the bulk of enterprises, but may well indicate a path for other corporations with large amounts of proprietary data, such as Moody’s, RELX/Elsevier and Morningstar.
What is more broadly true is that running open-weight models is a definite alternative to relying on the US frontier models like ChatGPT and Claude. Models like Kimi, Qwen, GLM and DeepSeek are close to the performance levels of US frontier models on many benchmarks, and actually exceed them in a few cases. Enterprises are catching on to this, with DeepSeek on its own now processing more tokens than OpenAI, Anthropic or Google. This change has been rapid, with open-weight models capturing 60% market share on the OpenRouter gateway (which Stripe agreed to acquire in August 2026 for a reported $7-8 billion) by June 2026, up from 40% market share in March 2026.
This seemingly inexorable rise in open-weight models poses interesting questions for the economics of the US frontier labs OpenAI and Anthropic, which are both seeking to go public in the near future at dizzying valuations, despite both currently haemorrhaging cash. OpenAI projected a $14 billion loss for 2026 after an operating loss of $25 billion in 2025. This projected loss may turn out to be optimistic based on the first half of 2026, with OpenAI’s Q2 2026 revenue up just 18% on Q1 2026 at $6.7 billion.
Investors may be prepared to pay up if they believe in a shining future of profitability for these companies, but open-weight models, which still incur infrastructure and operating costs but have no per-token model-provider fee, may present a large bump in the road to that future profitability.







