
Nearly two years after leaving OpenAI, former Chief Technology Officer Mira Murati has launched the first AI model from her new company, Thinking Machines Lab.
The model, called Inkling, is designed as an open-weight multimodal system that developers can download and customize. While the launch marks a major milestone for one of the most closely watched AI startups in Silicon Valley, one disclosure in the company’s technical report has drawn particular attention.
Thinking Machines Lab has confirmed that part of Inkling’s training process relied on synthetic data generated by Chinese open models. At a time when Washington is increasing scrutiny of Chinese AI technologies and tightening restrictions around advanced AI development, that decision is likely to become an important part of how policymakers evaluate the company going forward.
Inkling is Thinking Machines Lab’s first AI model
Thinking Machines Lab launched Inkling as its first in-house foundation model after raising billions of dollars in funding and assembling a team that includes several former OpenAI researchers.
The company describes Inkling as an open-weight mixture of experts model with 975 billion total parameters, although only about 41 billion are activated during any single task. It was trained on roughly 45 trillion tokens spanning text, images, audio, and video, giving it multimodal reasoning capabilities. At launch, however, the model generates text-based outputs, including code and structured data.
Unlike OpenAI’s GPT models or Anthropic’s Claude, Inkling allows developers to download and modify its weights, making it part of the growing open model ecosystem.
The company disclosed using Chinese AI-generated data
The technical report accompanying Inkling states that the model was trained using a mix of publicly available datasets and synthetic data generated by several existing AI models. Those included Chinese open models such as Moonshot AI’s Kimi family. Separately, Inkling’s mixture-of-experts architecture itself also draws on the design of DeepSeek-V3, the Chinese model from DeepSeek.
The company says synthetic data was used to improve reasoning and instruction following, an increasingly common practice across the AI industry. The report also notes that Thinking Machines filtered, evaluated, and combined outputs from multiple sources before incorporating them into training.
Although using outputs from open models is generally permitted under their licenses, the disclosure stands out because of the current geopolitical environment surrounding artificial intelligence, especially in the context of the United States and China.
Why the Disclosure Matters
The United States and China are locked in an increasingly intense AI competition. Washington has introduced export controls on advanced chips, imposed restrictions on Chinese technology companies, and raised concerns about national security risks linked to Chinese AI systems.
At the same time, Chinese open models have become increasingly competitive on performance and cost, with Moonshot AI’s Kimi K3 serving as a recent example. As such, several leading American AI executives have publicly warned that Chinese models are rapidly narrowing the gap with their U.S. counterparts.
Against that backdrop, a prominent American AI startup acknowledging that it benefited from synthetic data generated by Chinese models is likely to receive close attention from lawmakers and regulators, even though the practice itself is not unusual within the AI research community.
A Different Strategy from OpenAI
Thinking Machines Lab has repeatedly said it wants to build AI systems that are more customizable and transparent than today’s leading closed models. Inkling reflects that strategy through its open-weight release and emphasis on developer flexibility.
For Murati, the launch of Inkling serves as the first public demonstration of Thinking Machines Lab, the company she founded after leaving OpenAI.
Whether Inkling ultimately succeeds will depend on its technical performance and its continuous developer adoption. But its acknowledgment that part of its training pipeline relied on Chinese AI generated data ensures that discussions about the model will extend beyond benchmarks.
It also places Thinking Machines Lab at the center of a broader debate over how global AI development is becoming increasingly interconnected, even as governments move to separate their technology ecosystems.
