RobotAIGeek

World Models Explained: Why Liqing Intelligence Just Raised Hundreds of Millions to Build the 'iOS of Physical AI'

A three-month-old startup founded by a 28-year-old Tsinghua professor has closed one of 2026's most notable seed rounds in China's AI sector. Liqing Intelligence secured several hundred million yuan from an all-star investor roster including Sequoia China, Hillhouse Capital's GL Ventures, and Shunwei Capital. The company is building what it calls 'Physical AI Infrastructure,' a full-stack system combining proprietary data pipelines, a differentiable physics engine, and a native world model designed to train robots across any hardware form factor and any real-world scenario. This article explains what world models are, why they matter across industries from autonomous driving to home robotics, how Liqing's approach differs from global competitors including NVIDIA, Yann LeCun's AMI Labs, and Fei-Fei Li's World Labs, and what the funding surge signals for the broader physical AI ecosystem.

A
2 min readPosted: Jul 3, 2026
World Models Explained: Why Liqing Intelligence Just Raised Hundreds of Millions to Build the 'iOS of Physical AI'
What Is a World Model, and Why Does It Matter Now?

For the past several years, AI progress has been defined by large language models that predict the next word in a sequence. These systems learned from trillions of words of internet text and achieved remarkable capabilities in reasoning, code generation, and conversation. But they share a fundamental limitation: they learned how humans describe the world, not how the world actually works. A language model knows that the word 'gravity' appears near 'fall' and 'weight,' but it has no internal representation of what happens when you release a glass above a table.

A world model represents a paradigm shift from 'Next Token Prediction' to 'Next State Prediction.' Instead of predicting which word comes next, a world model predicts what happens next in a physical environment when an action is taken. In essence, a learnable physics simulator: an AI system that can imagine the consequences of actions before they are executed, understand spatial relationships and maintain a consistent internal representation of physical reality over time.

The urgency behind world models in 2026 stems from two converging pressures. First, the scaling laws that powered language model progress are showing diminishing returns. High-quality text data is approaching exhaustion and benchmark improvements per dollar of compute are flattening. Second, the embodied AI industry, which aims to deploy intelligent robots in factories, homes, and streets, has hit a data wall. Training a robot to perform physical tasks requires enormous volumes of real-world interaction data that is expensive, slow, and sometimes dangerous to collect. World models offer a solution: train robots in simulated physics, then transfer skills to reality with minimal real-world fine-tuning.

Where World Models Are Already Being Used
Autonomous Driving: The Most Mature Application

Autonomous driving represents the most commercially advanced use case for world models. Waymo published its World Model in February 2026, using it to simulate rare and dangerous driving scenarios at a scale impossible through road testing alone. NVIDIA's Cosmos platform, downloaded over two million times, generates synthetic driving data for companies including XPENG, Waabi, and Uber. In China, Huawei's autonomous driving division publicly rejected the VLA (Vision-Language-Action) approach in favour of its WA (World Action) architecture, while Geely launched a World Action Model that unifies smart driving, cockpit intelligence, and chassis control into a single physics-aware system. The efficiency gain is substantial: synthesizing rare dangerous scenarios through world models is estimated to be ten times more efficient than accumulating equivalent data through physical road testing.

Robotics and Embodied AI: Solving the Data Scarcity Problem

The central challenge in robot learning is data scarcity. Unlike language models that can train on the entire internet, robots must learn from physical interactions that happen in real time, in real space, with real objects. World models address this by allowing robots to 'practice' skills in simulated environments, then transfer those skills to physical hardware with minimal real-world calibration. NVIDIA's Cosmos 3, released in June 2026, currently ranks first across seven physical AI benchmarks covering world generation, robot action policy, and industrial vision. Leading humanoid and industrial robot companies including 1X, Agility Robotics, Figure AI, and Unitree are integrating world model training into their development pipelines.

Spatial Design, Architecture, and Smart Homes

Qunhe Technology, which listed on the Hong Kong Stock Exchange in April 2026 as the world's first 'spatial intelligence' public company, demonstrated that world models have commercial value beyond robotics. The company leverages over a decade of interior design data, comprising billions of physically correct room layouts, as training data for embodied AI systems. Its stock surged 144 percent on the first trading day. The application extends to smart home simulation, where world models can predict how a robot assistant would navigate a specific apartment layout, interact with furniture, and respond to household tasks before the robot is physically deployed.

Gaming, Film, and Virtual Worlds

Google DeepMind's Genie 3, released in late 2025, became the first real-time interactive world model capable of generating persistent 3D environments at 24 frames per second. Fei-Fei Li's World Labs launched Marble as a commercial product for generating explorable 3D worlds from text prompts, with pricing from free to $95 per month. Tencent's Hunyuan 3D World Model 2.0 compressed open-world game map creation from months of manual work to minutes of AI generation.

Liqing Intelligence: The Newcomer With a Full-Stack Thesis
Founding Story

Li Yiming was born in 1997. He completed his PhD in AI and Robotics at New York University in January 2025, where he published over 20 papers at top-tier conferences including CVPR, NeurIPS, and RSS, accumulating over 4,000 citations. His doctoral work focused on robust, efficient, and scalable computational models for 3D scene parsing and robot decision-making. During his PhD, he collaborated with Xie Saining, who would later become co-founder and Chief Scientist of Yann LeCun's AMI Labs.

After graduating, Li joined NVIDIA as a Research Scientist in the Vision and Robotics group, where he worked on foundation models and digital twins for autonomous robotics. He was one of only ten recipients globally of the NVIDIA Graduate Fellowship in 2024 and filed four US patents during his time at the company. In March 2026, he returned to China to join Tsinghua University's College of AI as an Assistant Professor. One month later, in April 2026, he founded Liqing Intelligence.

The Funding

Within two months of founding, Liqing Intelligence completed multiple funding rounds totalling several hundred million yuan in seed capital. The investor roster reads like a who's who of China's top-tier venture capital: Sequoia China, GL Ventures (Hillhouse Capital's early-stage arm), Shunwei Capital (Lei Jun's fund), Frees Fund, Starlink Capital, Shuimu Tsinghua Alumni Seed Fund, and SEE FUND. Notably, the round also includes strategic industrial investors: AgiBot (one of China's leading embodied AI companies), Linker Hand (a dexterous hand manufacturer), and Century Golden Resources. The presence of AgiBot as both investor and potential customer signals that Liqing's infrastructure may serve as a training backbone for existing robot hardware companies.

Technical Architecture: Physical AI Infrastructure

Li Yiming's core thesis is that a world model alone is insufficient. In a July 2026 interview with 36Kr, he used an analogy from the Tang Dynasty novel 'The Lychees of Chang'an': delivering fresh lychees from Guangdong to the capital required not just a fast horse, but an entire system of preservation, relay stations, route planning, and supply logistics. Similarly, he argues, a world model is 'just the horse,' worthless without the surrounding infrastructure of data collection, physics simulation, and deployment feedback loops.

Liqing's system comprises three self-developed components. The first is a proprietary data pipeline designed to scale collection from the industry average of hundreds of thousands of hours to millions or tens of millions of hours. The company has developed its own tactile gloves and wearable sensors that reduce per-unit data collection costs from dollar-level to yuan-level, enabling human workers in factories, hotels, and kitchens to generate training data at scale simply by performing their normal tasks.

The second component is a differentiable physics engine that implements a Real-to-Sim-Real closed loop. This engine can model complex materials including fluids, soft bodies, and objects undergoing elastic or plastic deformation, enabling robots to practice tasks like cutting, stirring, threading, and pouring in simulation before attempting them physically. The company claims that by aligning the state transitions in their physics engine with a small amount of real-world data, they can achieve equivalent task success rates using only one percent of the real robot data that competitors require.

The third component is the world model itself, which serves dual roles: as a self-supervised pre-training objective (modelling both state and action simultaneously) and as an interactive environment for reinforcement learning during post-training. The company's visual tokenizer, which converts physical world observations into machine-readable representations, reportedly already outperforms Meta's DINOv3 foundation model.

Global Competitive Landscape: Where Liqing Fits

The world model space in 2026 is crowded, well-funded, and globally distributed. To understand Liqing's position, it helps to map the field across three tiers based on funding scale, technical maturity, and commercial deployment.

uploaded image
Why Li Yiming Says Competitors Are Not Building 'Native' World Models

In one interview, Li Yiming offered pointed critiques of the three dominant approaches in the world model space, arguing that none of them qualifies as a 'native world model' capable of end-to-end physical intelligence.

On VLA (Vision-Language-Action) models, which companies like Google and several Chinese startups have adopted, Li argues that language is fundamentally a communication interface, not a modality for understanding the physical world. 'Language is a highly discretized space,' he said. 'Every country has different grammar rules, language is full of human biases, and many things simply cannot be expressed in words. Language exists for communication between humans, not as a native representation of physics.'

On Yann LeCun's JEPA architecture, which AMI Labs is commercializing with over $1 billion in funding, Li acknowledges its elegance but notes a critical limitation: 'JEPA can predict state, but it cannot output actions.' A system that understands what will happen but cannot decide what to do remains incomplete for robotics deployment.

On video generation models (including approaches derived from Sora, Runway, and several Chinese video AI companies), Li is equally direct: 'Video generation models can only fit the surface appearance of the world. They cannot guarantee the geometric and physical consistency required for complex task learning.' A generated video may look physically plausible to a human viewer while containing impossible physics that would cause a robot trained on it to fail catastrophically in the real world.

Liqing's alternative, which Li calls a 'native world model,' integrates perception, reasoning, decision-making, and action output into a single architecture trained end-to-end on physical interactions. Whether this claim holds up against benchmarks remains to be seen, as the company has not yet published comparative results.

What This Means for Robotics Adoption

World models represent the critical missing layer between AI intelligence and physical deployment. Without them, every robot must learn from scratch through expensive real-world trial and error. With them, a single training investment can theoretically scale across unlimited hardware platforms, environments, and tasks. The analogy to operating systems is deliberate: just as iOS allowed millions of apps to run on standardized hardware, a mature Physical AI Infrastructure could allow millions of robot skills to be developed, tested, and deployed without per-task physical data collection.

The funding concentration in this space during H1 2026 is extraordinary. China's world model startups alone have attracted over RMB 15 billion (approximately $2.1 billion) in the first half of the year, while globally the figure exceeds $4 billion when including AMI Labs, World Labs, and NVIDIA's Cosmos ecosystem investments. This capital intensity suggests that the industry consensus has shifted: world models are no longer a research curiosity but a prerequisite for scalable robot deployment.

For Liqing Intelligence specifically, the test will come at the end of 2026, when Li Yiming has committed to releasing a cross-scenario world model for enterprise customers. If the company's claim of achieving equivalent performance with one percent of competitors' real-world data holds in production environments, it would represent a genuine breakthrough in training efficiency. If not, it joins a long list of well-funded startups whose technical claims outpaced their delivery timelines. The 2028 milestone Li has set for scaled commercial deployment will be the definitive test.

Sources:

1. 36Kr Exclusive Interview with Li Yiming (July 1, 2026)

2. Tsinghua University College of AI, Faculty Profile: Yiming Li

3. Tech in Asia: Shunwei, Sequoia China back AI startup Liqing (July 1, 2026)

4. NYU Tandon School of Engineering: NVIDIA honors Ph.D. candidate Yiming Li (Dec 2023)

5. NVIDIA Research: Graduate Fellowship Profile

6. TMTPost: AI Newcomers Collectively Bet on World Models (June 23, 2026)

7. Securities Times: World Model Funding Feast (April 1, 2026)

8. NVIDIA Newsroom: Cosmos 3 Launch (June 2026)

9. Waymo Blog: The Waymo World Model (February 2026)


Hero Image: NVIDIA Cosmos World Foundation Model platform visualization. Source: NVIDIA Newsroom (publicly available press asset). Selected because NVIDIA Cosmos represents the world model infrastructure concept central to this article, and founder Li Yiming's direct professional connection to NVIDIA.

Disclaimer: This article is for informational purposes only and does not constitute investment advice. The information presented is based on publicly available sources and may not reflect the most current developments. Readers should conduct their own research before making any investment or business decisions.