LG Targets 100000 Hours of Humanoid Training Data
LG Electronics told NVIDIA executives in Seoul that its Yangjae data factory should produce 100,000 hours of robot training data by year end, pairing CLOiD robots with NVIDIA physical AI tools.

LG Electronics hosted NVIDIA robotics executives at its Yangjae research and development campus in Seoul on Tuesday and said the data factory under construction there is expected to generate 100,000 hours of robot training data by the end of 2026. The visit followed a 13 August memorandum of understanding in Santa Clara, California, and put a number on how fast LG wants to turn household, factory, and logistics work into a reusable physical artificial intelligence stack.
Madison Huang, NVIDIA senior director of product marketing for Omniverse and robotics, toured the site with LG Electronics Chief Executive Officer Lyu Jae-cheol, LG CNS Chief Executive Officer Hyun Shin-gyoon, and LG Sciencepark President Chung Sue-hyun. LG CLOiD home robots are already collecting task data in a mock home, a replica of LG’s Tennessee washing machine plant, LG CNS logistics cells, and LG Innotek robotic-hand stations. NVIDIA Omniverse libraries, Cosmos world models, and the Isaac robotics platform are being used to augment, synthesize, and replay that data.
For a procurement or operations lead, the useful fact is not the bouquet that CLOiD handed Huang. It is that a large appliance and electronics group is treating training data as a capital project with a floor area, a robot count, a year-end operating date, and a named software stack. LG said the Yangjae building covers about 10,000 square meters across one basement level and three floors above ground, should house several hundred robots by year end, and is scheduled for full operation before 2027. The 100,000-hour figure includes both hours collected on site and hours generated synthetically. LG compared that volume with roughly 12 years of continuous operation.
That comparison is a marketing conversion, not a physics claim. Twelve years of one robot running without pause is not the same as 100,000 hours of diverse, labeled, safety-checked episodes. Buyers should treat the hour count as a capacity target and ask what share is real-robot time, what share is synthetic, which tasks are represented, and how failure cases are labeled.
Seoul Turns a Showroom Into a Data Plant
LG Group and NVIDIA signed the Santa Clara memorandum four days before the Seoul walkthrough. The companies framed the pact around physical AI, AI infrastructure, and mobility. The speed of the follow-up visit is itself a signal. Large Korean groups often announce platform alliances and then spend months translating them into plant-level work. Here the first public inspection happened inside a week, with the robotics data factory already under construction and CLOiD units already generating scenes.
LG has also built an organizational container for the spend. Last month it created a Robotics Business Center that reports to the chief executive. The center is supposed to pull industrial, commercial, and home robots into one execution unit rather than leaving them as scattered product lines. LG already sells or pilots robots in factories and commercial sites and said it now wants a home-robot line as well. The wider group also has capabilities in actuators, sensors, and batteries through LG Electronics, LG Innotek, and LG Energy Solution.
The Tuesday tour showed how those pieces are being staged as a closed loop. In the home replica, CLOiD units repeat cleaning tasks. In the Tennessee replica, they move, stack, and assemble parts. In the logistics and hand-training rooms they collect data on the more specialized motions that warehouse software and gripper programs actually need. NVIDIA tools then expand those traces. LG’s half-year report, released Friday, said joint projects already cover manufacturing robots from proof of concept through deployment on production sites.
None of that yet proves a humanoid can run a shift. It does prove LG is no longer treating humanoid work as a concept video. The company said the data will feed a robot foundation model intended to help machines perceive surroundings, understand instructions, and carry out physical tasks. Humanoids, LG noted, have to handle many objects and changing rooms rather than repeat a short programmed cycle.
Hours Are Cheap Until Labels, Safety, and Replay Are Priced
The hard commercial question is whether 100,000 hours is a training corpus or a storage bill. Physical AI teams already know that uncurated video is expensive to keep and cheap to waste. A useful hour is one that can be replayed, segmented, associated with force or pose signals, and used to train a policy that transfers to a different room. A useless hour is a robot wandering a mock kitchen while the interesting collision happens off camera.
LG’s design tries to reduce that waste by copying real plants and homes instead of filming random demos. The Tennessee cell matters because it is tied to a factory LG already runs. If the same parts, fixtures, and cycle times appear in Seoul and in the United States plant, synthetic augmentation has a real geometry to cling to. If the replica is only loosely similar, Cosmos-generated scenes will look rich and still fail on the first unexpected tote.
Buyers should also separate three products that LG is bundling in one tour. The first is a data factory: a building, robots, labeling pipeline, and simulation stack. The second is a robot foundation model: software that may or may not be licensed outside the group. The third category is finished robots and components: CLOiD, actuators, hands, and later a bipedal humanoid that LG has separately said it wants to show in the first quarter of 2027. A plant that needs a pallet robot next year does not buy a 2027 humanoid announcement. It might buy a data-collection method, a simulation seat, or a component quote.
NVIDIA’s role is the other price driver. Omniverse, Cosmos, and Isaac are not free utilities. They are a vendor stack that can lock a factory’s digital twin, synthetic data, and edge runtime to one graphics and software supplier. That can be rational. It can also create switching costs that dwarf the robot hardware invoice. Operations teams should ask which data stays on LG premises, which models are trained on NVIDIA cloud or on-prem graphics processing units, and whether a third-party integrator can replay the same corpus on another stack.
The public-interest version of the story is a robot handing flowers to an executive. The procurement version is more prosaic. Who owns the traces from a Tennessee line once they have been synthesized in Seoul? If a contract manufacturer later uses a different robot brand, can the hours transfer? If a safety incident occurs, can the company reconstruct the exact training distribution that produced the policy? Those questions decide whether a data factory is an asset or a liability.
What a Buyer Should Demand Before Copying the Model
LG is not the first company to say that physical AI is a data problem. It is one of the first large appliance manufacturers to publish a year-end hour target, a floor plate, and a robot headcount for a dedicated factory whose output is training data rather than washers. That is why the story matters outside Korea. Chinese humanoid vendors have been racing to collect factory hours. United States and European teams have been racing to collect warehouse hours. LG is trying to collect both household and plant hours inside one group that already ships products into those rooms.
The honest limit is that Tuesday’s event was a supervised visit, not an independent audit. LG did not publish a breakdown of real versus synthetic hours, a list of task taxonomies, a safety case, or a customer who has bought the foundation model. The 100,000-hour target is a 2026 year-end claim. Several hundred robots in a 10,000-square-meter building is a capacity plan. Full operation by year end is a construction and staffing plan. Any of those can slip without the memorandum of understanding being withdrawn.
A practical checklist still follows from the disclosed design. Ask for the share of hours that come from real robots versus generated scenes. Ask which plants, besides the Tennessee replica, will contribute live data. Ask whether CLOiD traces are being collected under a privacy and labor protocol that would survive a deployment in Europe or the United States. Ask how LG Innotek hand data and LG CNS logistics data are joined into one model, or whether they remain separate specialists. Ask what “total robotics solutions provider” means in a request for proposal: a robot, a cell, a software subscription, or a multi-year data service.
The bigger shift is that robot buying is starting to look like semiconductor buying. The visible machine is not the scarce item. The scarce item is a closed loop that can turn yesterday’s shift into tomorrow’s policy without sending a process engineer back to the teach pendant. LG is spending against that loop. NVIDIA is supplying the simulation and synthesis layer. The test will be whether the Tennessee line, not the Seoul tour, runs longer with fewer interventions.
Lyu said LG would use group capabilities and global partners to compete in physical AI. That sentence is a strategy, not a shipment. The shipment to watch is more boring and more important: how many labeled hours leave Yangjae in a form a factory team can use, and whether those hours cut integration time on a real line.
-----
This article is for informational purposes only and does not constitute investment, legal, engineering, or procurement advice. Buyers should verify current specifications, commercial terms, safety certification, and regulatory status with the relevant companies and authorities before making purchasing or partnership decisions. Analysis synthesizes company statements and public market activity.











