Field guide
Physical AI in Production: Six Deployments Compared by Evidence
A source-backed comparison of Gritt, Sereact, Monumental, Digit, Figure 02, and AGIBOT G2 across task, customer confirmation, fleet, runtime, throughput, intervention, safety, and commercial access.
The short answer
Physical AI is already doing recurring work, but the best public evidence comes from narrow workflows, not general capability. Construction systems handle solar panels and bricks. Warehouse systems pick items or move totes. Industrial robots load automotive parts and run tablet-inspection stations. The physical forms differ, but the evidence questions are the same.
This comparison applies one framework to six cases:
- Gritt, robotic arms mounted on existing construction equipment for solar installation;
- Sereact Cortex, AI software operating across warehouse picking systems;
- Monumental, a mobile bricklaying system sold as a subcontracted construction service;
- Agility Robotics Digit, a bipedal humanoid moving totes at GXO;
- Figure 02, a bipedal humanoid that loaded sheet-metal parts at BMW;
- AGIBOT G2, a wheeled dual-arm robot serving a tablet quality-inspection line at Longcheer.
The evidence does not support a simple conclusion that specialized machines work while humanoids do not. Digit, Figure 02, and G2 all crossed into bounded production work. Nor does it prove that a large cumulative output equals high uptime or good economics. Most sources omit fleet size by site, interventions, maintenance, incidents, and commercial terms.
The useful conclusion is more precise: a well-scoped task makes performance easier to measure, regardless of robot shape. Production claims become defensible when a named customer, defined workflow, operating duration, throughput denominator, and human-assistance record appear together. No case reviewed here publishes all of them.
Evidence at a glance
Status reflects primary and attributable reporting reviewed on July 22, 2026.
| System | Workflow | Strongest published evidence | Customer confirmation | Main missing evidence |
|---|---|---|---|---|
| Gritt | Solar-panel pick, carry, and placement | Two field systems reported by TechCrunch; Gritt reports 30,000+ panels and 18+ MW installed | TechCrunch spoke with one unnamed customer | Site names, duration, uptime, intervention, audited safety and productivity |
| Sereact Cortex | Warehouse picking and returns across robot arms and cells | Sereact reports 200+ live systems and 1B+ production picks | Zalando is a named strategic investor; the cited release does not attribute the totals to customer records | Per-site fleet, time window, errors, interventions, maintenance, independent audit |
| Monumental | Bricklaying delivered as a subcontracted service | Monumental reports 100+ robots on active European job sites | Company describes completed projects, but the cited funding release is the source for the fleet claim | Utilization, robot mix, runtime, intervention, comparable crew baseline, incidents |
| Digit | Moving totes between work areas at GXO | Agility reports 100,000+ totes in live commercial deployment | GXO confirms a multi-year robotics-as-a-service agreement | Fleet size, runtime, rate, interventions, uptime, economics |
| Figure 02 | Loading sheet-metal parts for BMW X3 production | Figure reports 1,250+ runtime hours and 90,000+ parts; BMW confirms 10 months and 30,000+ vehicles supported | Named customer with operating duration and production context | Robot count, interventions, cycle success, shift uptime, safety record, commercial terms |
| AGIBOT G2 | Loading and unloading at tablet quality-inspection stations | Longcheer confirms an 8-hour no-intervention test; AGIBOT reports 64+ cumulative hours and about 3,000 units per shift | Named customer with one bounded no-intervention test | Fleet size, shift length, full-period interventions, maintenance, audit, economics |
A production-evidence framework
A video can establish that a behavior occurred. A customer announcement can establish that a relationship exists. A production claim needs more.
This article uses eight fields:
- Named customer or operator. Who owns the workflow and can confirm the result?
- Bounded task. What enters the process, what action occurs, and what output counts as complete?
- Fleet size. How many robots produced the reported total at each site?
- Operating duration. Were the robots active for one test, several shifts, months, or years?
- Throughput. Is output reported as attempts, successful units, parts, picks, totes, or another measurable unit?
- Human involvement. How often did people teleoperate, intervene, reset, clear faults, or recover the system?
- Safety and maintenance. What incidents, damage, stoppages, component replacements, and scheduled service occurred?
- Commercial access. Was this a paid service, contract, lease, robotics-as-a-service agreement, pilot, or internal project?
These fields complement the announcement-to-deployment labels in our humanoid robot availability tracker. They also matter for fixed arms, mobile construction equipment, and other forms of physical AI.
Gritt: two field systems, strong output claims, limited operating detail
Gritt launched publicly on July 21 with $32.4 million in pre-seed and Series A funding. Our physical AI funding tracker compares that capital with six other companies by the milestones it has produced. Gritt’s system mounts robotic arms and sensing on existing construction equipment such as skid steers and forklifts rather than introducing a new general-purpose robot body (Gritt).
The initial production task is bounded and legible. The machine unloads large solar panels, carries them to tracker structures, and positions them while workers complete fastening. Gritt’s solar page reports:
- more than 30,000 panels installed;
- more than 18 MW of solar installed;
- 2.8 GW contracted;
- a claimed fourfold productivity increase reported by engineering, procurement, and construction customers;
- claimed reductions in cost and heavy-load injuries (Gritt solar).
Those figures are not equal evidence. Installed panels and capacity describe completed output. Contracted capacity describes future work. Productivity, cost, injury reduction, zero breakage, and safety language are company claims without a published protocol, comparison window, incident denominator, or outside audit.
TechCrunch adds useful independent reporting. It says two systems were deployed in the field, describes the hardware configuration, and reports that it spoke with one customer who declined to be named. The article attributes the projected increase from 800 panels per day to 3,000 to 4,000 panels per day to Gritt’s CEO, not to an audited customer study (TechCrunch).
Evidence classification: field deployed for a bounded construction workflow. Completed output is company-reported, and an independent publication confirms two field systems plus one unnamed customer conversation. Recurring duration, site-level fleet logs, interventions, downtime, maintenance, incidents, and normalized productivity remain undisclosed.
Sereact: the largest cumulative output claim, with weak denominator detail
Sereact supplies AI software for warehouse robots rather than one fixed body. Its Cortex platform runs across single-arm picking cells, dual-arm returns stations, and other configurations. That makes it a useful case for separating the intelligence layer from robot morphology.
In a July 20 release announcing Zalando as a strategic investor, Sereact reports:
- more than 200 systems live across Europe;
- more than one billion real production picks completed on Cortex;
- named customers including Daimler Truck, Mercedes-Benz, BMW, MS Direct, Active Ants, DeltiLog, Rohlik Group, and Austrian Post (Sereact).
One billion picks is the largest cumulative task count in this comparison. It is still difficult to normalize. The release does not define the measurement period, allocate picks across customers or hardware, state whether the count includes retries, or publish failures, human corrections, system-hours, and maintenance. More than 200 live systems also does not reveal how many are active during a typical shift or how deployment age varies.
Zalando’s investment is a named operator-side commercial signal, supported by a quotation from a Zalando executive in Sereact’s release. It does not independently verify the 200-system or one-billion-pick totals. The distinction matters because strategic confidence and operating verification answer different questions.
Evidence classification: company-reported recurring production operation across a large installed base, with named customers and a strategic operator investor. The cumulative figures are meaningful but lack customer-side breakdowns, error and intervention denominators, and independent audit.
Monumental: more than 100 robots, but fleet composition matters
Monumental uses small autonomous mobile machines to transport bricks and mortar, then places bricks with a dedicated arm. It offers the work as a subcontractor, which means customers buy completed masonry scope rather than purchasing and integrating a robot.
In its July Series B announcement, Monumental says it now runs more than 100 robots laying brick on real job sites across Europe (Monumental). That is a substantial field-fleet claim and a direct counterexample to the idea that physical AI must be humanoid.
The number needs a morphology and accounting caveat. A bricklaying system can include multiple mobile machines with distinct transport, mortar, and placement roles. A count of robots is not necessarily a count of complete bricklaying crews or simultaneously active systems. The release does not publish fleet utilization, runtime, bricks per operating hour, intervention frequency, weather-related stoppages, rework, maintenance, incidents, or a matched manual baseline.
The subcontracting model is commercially important. It can reduce adoption friction because a contractor does not need to buy a robot, train an internal robotics team, or redesign procurement around a capital asset. It also makes public economics harder to inspect because project pricing, labor mix, margin, and service obligations remain private.
Evidence classification: company-reported recurring field operation and commercial service with more than 100 robots. Project delivery is established at a high level, while normalized fleet utilization, throughput, assistance, safety, and economics remain undisclosed.
Digit at GXO: customer-confirmed commercial access, vendor-reported throughput
Digit is a bipedal humanoid, but its production task is as narrow as the construction and warehouse cases above. At GXO’s Flowery Branch operation, Digit moves totes between work areas and integrates with existing automation.
GXO announced a multi-year robotics-as-a-service agreement with Agility Robotics in June 2024. That customer source establishes commercial access and operational intent, not just a lab partnership (GXO).
Agility later reported that Digit had moved more than 100,000 totes in live commercial deployment (Agility Robotics). The count is a concrete output measure. It does not reveal the number of robots, total scheduled hours, totes per robot-hour, rejected moves, human interventions, fault recovery, uptime, or the service cost paid by GXO.
This is production evidence, but not evidence of general-purpose warehouse autonomy. The robot performs a bounded tote workflow inside a prepared operation. The bipedal form may help it use spaces and equipment designed for people, but the public record does not provide a controlled comparison against a wheeled mobile robot, conveyor change, or other automation path.
Evidence classification: customer-confirmed multi-year commercial deployment with vendor-reported cumulative throughput. Fleet, runtime, intervention, reliability, safety, and unit economics remain incomplete.
Figure 02 at BMW: the strongest combined customer and vendor record
Figure 02 worked in BMW Plant Spartanburg’s body shop, loading sheet-metal parts into fixtures for welding. The task required repeated part handling on an active automotive line.
Figure reports that the program:
- ran a 10-hour shift Monday through Friday;
- loaded more than 90,000 parts;
- accumulated more than 1,250 hours of runtime;
- contributed to production of more than 30,000 vehicles (Figure).
BMW separately confirms that Figure 02 supported production of more than 30,000 BMW X3 vehicles over ten months. BMW calls the completed work both a pilot and a successful deployment, and says the robot performed precise, repeatable steps under real production conditions (BMW Group).
This is the strongest evidence chain in the comparison because the customer confirms the task, production setting, duration, and vehicle context, while the vendor supplies runtime and part counts. The sources still omit robot count, task attempts, rejected placements, human interventions, teleoperation, cycle-time distribution, scheduled versus actual uptime, maintenance, incidents, and commercial terms.
Generation boundaries also matter. Figure 02 was retired. Figure 03 has demonstrated a different sequencing workflow at BMW, but it has not inherited Figure 02’s operating record. A previous generation’s production result can guide expectations, not prove the successor’s reliability.
Evidence classification: customer-confirmed recurring production deployment for a bounded automotive workflow, supported by vendor runtime and throughput. It is not evidence of broad autonomy, and key reliability and assistance metrics remain private.
AGIBOT G2 at Longcheer: one clean intervention statement, then a broader vendor record
AGIBOT G2 is a wheeled, dual-arm industrial robot. At Longcheer’s Nanchang facility, it loads and unloads tablet quality-inspection equipment, places material, interfaces with stations, and returns finished units.
Longcheer provides unusually specific customer-side evidence. It says G2 was deployed on a tablet mass-production line and completed an uninterrupted eight-hour test without human intervention at a tablet precision-testing station on April 14 (Longcheer Technology). That statement is valuable because it binds autonomy to a task and a duration.
AGIBOT later published a broader record. After a six-day factory livestream, it reported more than 64 hours of operation across multiple robots (AGIBOT). Its July update says integration took 36 hours, the workflow supported approximately 3,000 units per shift, and downtime loss remained below 4% (AGIBOT).
The metrics should not be merged carelessly. Longcheer’s no-intervention statement covers one eight-hour test, not the full 64-hour record. AGIBOT does not define shift length or downtime-loss methodology. It also does not disclose fleet size, interventions, failures, scheduled maintenance, product defects, or whether the reported throughput counts inspected units, successfully completed units, or all line units handled during the shift.
Evidence classification: customer-confirmed production-line placement and one bounded no-intervention test, plus vendor-reported integration, runtime, throughput, and downtime. Longer-term reliability, fleet normalization, maintenance, and economics remain undisclosed.
What this comparison shows
1. Task scope is more informative than robot shape
The cases span mounted arms, warehouse cells, fleets of small construction robots, a bipedal logistics robot, a bipedal manufacturing robot, and a wheeled dual-arm industrial robot. All become measurable only when attached to a specific workflow.
A humanoid can have strong evidence when the task and customer record are clear. Figure 02 is the best example here. A specialized system can still have weak evidence if only aggregate marketing totals are available. Morphology affects engineering and workflow fit, but it does not determine evidence quality.
2. Cumulative output needs time and fleet denominators
One billion picks, 100,000 totes, 90,000 parts, and 30,000 panels sound directly comparable. They are not.
A useful rate needs at least successful output, system-hours, and fleet size. A useful reliability measure also needs attempts, interventions, resets, downtime, and maintenance. Without those denominators, cumulative output establishes scale of activity but not uptime, labor displacement, or return on investment.
3. Customer confirmation changes the evidence class
BMW, GXO, and Longcheer each confirm part of the operating record. Gritt has one unnamed customer conversation reported by TechCrunch. Sereact and Monumental identify commercial activity, but the headline operating totals in the reviewed releases remain company-reported.
Customer evidence is not automatically independent. Partners can share incentives and public-relations language. It is still stronger than a vendor claim alone because it establishes that the workflow exists inside the operator’s environment.
4. Intervention is the most important missing metric
Longcheer’s eight-hour statement is the only explicit, bounded no-intervention result in this set. It should not be generalized beyond that test.
For every other case, public sources do not quantify how often people teleoperated, recovered, reset, cleared materials, changed task parameters, or performed maintenance. Those actions may be entirely reasonable in production. Hiding them prevents readers from estimating staffing and reliability.
5. Commercial models shape what adoption means
GXO uses a robotics-as-a-service agreement. Monumental acts as a subcontractor. Sereact supplies software across multiple hardware forms. Gritt attaches to existing construction equipment. BMW and Figure describe a direct manufacturing program.
A buyer comparing these systems should ask what outcome is purchased, who owns the hardware, who operates it, what uptime is guaranteed, how exceptions are staffed, and how performance risk is allocated. A robot’s purchase price alone cannot answer those questions.
Agility has since published a three-stage Customer Acceleration Program that ends with a stated 90-day production measurement window before ongoing RaaS. Our humanoid Robots-as-a-Service guide separates that deployment method from the customer-level uptime, intervention, support, and ROI results that remain unpublished.
What buyers and operators should request
Before calling a physical-AI system production-ready, ask for a site-specific evidence pack:
- Workflow definition: inputs, outputs, task boundaries, excluded cases, and success criteria.
- Fleet and version record: robot count, hardware generation, software checkpoint, and configuration changes.
- Operating schedule: planned hours, actual active hours, shift structure, and deployment duration.
- Trial accounting: attempts, completed tasks, partial completions, defects, retries, and excluded runs.
- Human support: teleoperation minutes, interventions, resets, recoveries, supervision, and on-site staffing.
- Reliability: uptime, mean time between failures, maintenance time, component replacement, and spare requirements.
- Safety: hazards, incidents, near misses, stoppages, applicable standards, and site risk controls.
- Baseline: the manual or automated process used for cost, speed, quality, and safety comparisons.
- Commercial terms: contract structure, service levels, support, integration cost, performance guarantees, and exit conditions.
- Customer confirmation: a statement from the operator that distinguishes a test, pilot, paid service, and recurring operation.
For scope questions involving industrial, service, and mobile robots, use our robot safety standards guide. Safety claims require a defined application and risk assessment, not only a company adjective.
Evidence watchlist
The next disclosures that would materially improve this comparison are:
- Gritt site duration, interventions, audited productivity, breakage, and incident denominators;
- Sereact customer-side confirmation of system count, pick totals, error rates, and interventions;
- Monumental fleet composition, robot-hours, bricks per active hour, rework, weather downtime, and project-level customer results;
- Digit fleet size, system-hours, intervention rate, uptime, and service economics at GXO;
- Figure 02 trial and intervention logs, plus equivalent recurring metrics for the new Figure 03 workflow;
- G2 fleet size, shift definition, full 64-hour intervention record, maintenance, defect handling, and longer operating history.
Verdict
Physical AI has crossed into real construction, warehouse, logistics, and manufacturing work. The deployments are useful precisely because they are bounded. Each robot or software system has a defined material flow, environment, and operator.
The strongest case in this set is Figure 02 at BMW because customer confirmation, operating duration, runtime, and throughput appear across two primary sources. G2 at Longcheer has the clearest bounded no-intervention statement. Digit has a customer-confirmed multi-year commercial model and concrete vendor throughput. Gritt, Sereact, and Monumental show that task-specific systems can accumulate field and production evidence without adopting a humanoid body.
None publishes the full record needed to compare reliability or economics cleanly. The next step for the industry is not a larger headline number. It is a denominator: robots, hours, attempts, interventions, downtime, incidents, and cost for one clearly defined workflow.
Frequently asked questions
Is physical AI already used in production?
Yes, in bounded workflows. The reviewed evidence includes solar-panel handling, warehouse picking, bricklaying, tote movement, automotive part loading, and tablet inspection. The quality of public evidence varies sharply, and none of these cases establishes a robot that can perform arbitrary work.
Do task-specific robots have stronger deployment evidence than humanoids?
Some do, but robot shape is not the deciding factor. Gritt, Sereact, and Monumental report field or production operation, while Digit, Figure 02, and AGIBOT G2 also have meaningful production evidence. Task scope, workflow integration, customer confirmation, runtime, throughput, and intervention disclosure matter more than morphology.
Which case has the strongest customer confirmation?
BMW provides the clearest customer-side operating summary in this set, confirming that Figure 02 supported production of more than 30,000 BMW X3 vehicles over ten months. GXO confirms a multi-year commercial Digit agreement, while Longcheer confirms G2 deployment and an eight-hour test without human intervention. Other totals rely more heavily on vendor reporting.
What does more than one billion robot picks prove about Sereact?
Sereact reports more than one billion production picks on Cortex across more than 200 live systems. That is a substantial first-party cumulative operating claim, but the cited release does not provide a customer-by-customer breakdown, time window, error rate, intervention rate, or independent audit.
What deployment metrics are still missing most often?
Fleet size by site, scheduled operating time, successful throughput, human interventions, resets, recoveries, maintenance, safety incidents, software version history, and commercial terms are the most common gaps. A cumulative output number without these fields cannot establish uptime or unit economics.
Does a production deployment prove general-purpose autonomy?
No. Every case in this comparison performs a bounded workflow. Production evidence can establish useful, recurring work under specified conditions, but it does not transfer automatically to unrelated tasks, sites, hardware, or operating constraints.