Field guide
Sunday Robotics ACT-2: What 778 of 785 Laundry Folds Actually Prove
A source-backed audit of ACT-2’s reported 99.1% laundry-folding result, including scope, adaptation cost, grading, speed, generalization claims, and the evidence still needed from home beta testing.
The short answer
Sunday Robotics has published one of the most detailed first-party reliability reports yet for a home-manipulation robot. In its July 17, 2026 ACT-2 preview, Sunday reports 778 successful autonomous laundry folds in 785 attempts, an overall success rate of 99.1% ± 0.3% standard error. The company also declares the task boundary, uses one fixed model checkpoint across unseen homes, states that no home- or garment-specific deployment adaptation was used, publishes category-level denominators, defines a fold-quality rubric, describes two-stage human grading, reports completion time, and links the evaluation videos (Sunday Robotics).
That disclosure is substantially stronger than a highlight reel or an isolated percentage. A reader can examine what counted as success, which garments were included, where performance was weakest, how long successful attempts took, and how much adaptation was required at each home.
It is not independent replication, a safety case, or proof of a generally capable home robot. Sunday ran and graded the evaluation. The article does not quantify remote intervention, manual resets outside autonomous recovery, safety stops, excluded or aborted attempts, total operating hours, household-level sample sizes, or failures during broader unsupervised use. Memo’s family beta is planned for fall 2026; it is not a completed recurring deployment.
The right conclusion is narrow but important: ACT-2 provides unusually strong first-party evidence that one fixed robot policy can fold a declared range of laundry with high measured reliability across the company’s unseen-home evaluation. The beta must now show whether that result survives field operations, safety constraints, and real household usage.
Evidence at a glance
| Question | Published evidence | Classification |
|---|---|---|
| Was there a denominator? | 778 successes in 785 autonomous attempts | Strong first-party disclosure |
| Was task scope declared? | Nine garment categories, sizes XXS–8XL, varied materials, scenes, surfaces, lighting, orientations, and starting configurations | Bounded, documented scope |
| Were test homes used for training? | Sunday says no evaluation-home or physical-garment data was used for task-specific post-training or model selection | Zero per-home adaptation claimed |
| Was one model used? | Same checkpoint and system configuration across the reported evaluation | Fixed-policy evaluation claimed |
| Was quality measured? | Mean 4.72/5 across 778 completed folds; 98.3% rated four or five stars | Company-designed, human-graded rubric |
| Was speed reported? | Median 2:13 and mean 2:19 for successful folds, including autonomous retries and recovery | Useful successful-attempt timing |
| Are videos available? | Sunday links a complete evaluation-video compilation | Inspectable first-party evidence |
| Was it independently replicated? | No independent reproduction identified | Not established |
| Is it deployed in homes? | Family beta planned for fall 2026 | Scheduled pilot, not current deployment |
| Are interventions and safety stops quantified? | Not reported in the preview | Critical field evidence missing |
What Sunday actually evaluated
Laundry folding can sound like a single task, but the declared boundary matters. Sunday includes T-shirts, thick and thin long-sleeved tops, polos, sleeveless tops, blouses, pants, leggings, and shorts. The company says garments varied from XXS to 8XL across colors, materials, thicknesses, and textures. Socks, bras, underwear, and accessories were excluded because their ordinary treatment may involve pairing, sorting, or hanging rather than the same folding process (ACT-2 preview).
The scene scope includes unseen rooms, varied beds and folding surfaces, different lighting, and robot positions at either side or the foot of a bed. Garments began in baskets, piles, on beds, or on the ground, with arbitrary orientation and natural crumpling. Sunday says no evaluation-home data, target garments, expert demonstrations in those homes, or home-specific fine-tuning were used. The model weights remained fixed throughout the reported evaluation.
That is a meaningful zero per-home adaptation claim. It should not be expanded into “zero task-specific data.” Sunday says ACT-2 uses a large human pretraining dataset and a post-training loop with in-house Memo demonstrations. The key claim is that these improvements transferred to held-out homes and garments without collecting new task-specific data at each deployment site.
Reconstructing the 99.1% result
Sunday defines a successful attempt as autonomously folding and stacking the garment. Across the 785 attempts, seven failed and 778 succeeded.
The category table is especially useful because it shows unequal sample sizes:
| Garment type | Attempts | Reported success |
|---|---|---|
| T-shirts | 312 | 99.0% |
| Shorts | 98 | 100% |
| Thick long-sleeved tops | 85 | 100% |
| Pants | 85 | 98.8% |
| Thin long-sleeved tops | 79 | 100% |
| Leggings | 54 | 96.3% |
| Polos | 46 | 100% |
| Blouses | 19 | 94.7% |
| Sleeveless tops | 7 | 100% |
| Overall | 785 | 99.1% |
A 100% result on seven sleeveless tops carries much less information than 99.0% across 312 T-shirts. The overall result is therefore more informative than treating every category percentage as equally mature. Blouses were the lowest-performing category at 94.7%, and Sunday explicitly labels its explanation—that lightweight, deformable construction offers fewer stable geometric cues—as a hypothesis rather than a demonstrated cause.
Sunday also breaks performance down by initial configuration and robot position. It reports 98.8% for a pile on a bed across 514 attempts, 100% for a basket on a bed across 73, and 99.5% for a basket on the ground across 198. Results were 98.7% from the left side of the bed, 100% from the right, and 98.5% from the foot. Those cuts help test variation, although they are not necessarily statistically independent: the same attempt can belong to a garment, starting-position, robot-position, and sheet-color category.
Quality was measured separately from task completion
A garment can be technically folded while still being messy or unstable. Sunday grades every completed fold with a five-star rubric. A fold starts at five stars, with one star deducted for each flaw category:
- more than two inches of material overfolded inward;
- corresponding edges misaligned by more than two inches;
- a sleeve, leg, or hood extending more than two inches beyond the fold;
- a stacking error that creates excessive overhang or disturbs the fold.
Sunday reports a mean of 4.72/5 across 778 completed folds. It says 98.3% received four or five stars and 73.8% received five stars. The company says one annotator assigned the initial grade, a second independently reviewed it, differences were resolved by a review lead, annotators trained on reference folds, and the rubric was frozen before the evaluation.
This is better than leaving “success” undefined. It is still a company-designed and company-administered rubric. The preview does not report inter-rater agreement, the number of annotators, whether graders were blinded to model version or experimental condition, or a matched human baseline over the full sample. Side-by-side human examples are illustrative, not a controlled comparative study.
The speed result: useful, with one denominator caveat
For the 778 successful folds, Sunday reports a 2 minute 13 second median and 2 minute 19 second mean. Timing begins when Memo starts retrieving a garment and ends when it adds the folded garment to the stack. Autonomous retries and recovery are included.
This provides a much clearer operating measure than a sped-up video. But it is a successful-attempt time distribution. The article does not provide time-to-failure for the seven unsuccessful attempts, setup time between attempts, battery and charging effects, intervention time, or sustained throughput over a complete household laundry load. It therefore supports comparison of successful fold duration, not a full labor, utilization, or household-productivity model.
Generalization and one-example learning claims
Sunday’s broader research claim is that scaling pretraining narrows the gap between in-domain and out-of-domain performance after post-training. The company says ACT-2 is pretrained on a high-quality, diverse sensorized human dataset and then improved with curated in-house Memo data. Its charts report that larger pretraining scales reduce the post-training generalization gap, while higher-quality data improves validation loss and downstream success at matched data volume and compute (ACT-2 preview).
The article also reports four one-example experiments. Sunday post-trained four independent copies of the same pretrained model with one demonstration of a different folding technique, then says each checkpoint reproduced its newly learned technique on a held-out garment.
These experiments are promising mechanism evidence, but the public article does not disclose enough to treat them as a general one-shot-learning benchmark:
- model architecture, parameter count, and complete training recipe are not published;
- the pretraining dataset’s size and composition are not quantified publicly;
- exact in-domain and out-of-domain environment counts are incomplete;
- the one-example test uses four behaviors and a held-out garment rather than a broad task suite;
- model checkpoints, code, data, and an independent reproduction are unavailable;
- the 99.1% laundry result and the four one-example behavior tests answer different questions and should not be merged.
The responsible reading is that Sunday presents first-party experimental evidence for efficient post-training transfer in its own system. It does not yet establish that any arbitrary new robot behavior can be learned from one demonstration.
Sunday’s “Solve” framework is useful—even without adopting the label
Sunday proposes a three-part reporting framework called a Solve:
- Performance: success, quality, and speed within a stated boundary;
- Scope: the objects, environments, and configurations covered by the claim;
- Adaptation cost: new data, demonstrations, fine-tuning, intervention, or system modification required for each deployment.
The underlying idea is sound and portable. A reliability percentage without its task distribution and adaptation requirement is difficult to compare. Ninety-nine percent on one prepared garment in one room is a different result from ninety-nine percent across hundreds of varied garments in held-out homes with no site-specific fine-tuning.
For publication and buyer diligence, we would add four more fields: independence, intervention, duration, and safety. Those distinguish a company-run evaluation from external replication, autonomous recovery from human rescue, hundreds of selected attempts from sustained field operation, and functional success from safe use around household members.
| Evidence dimension | ACT-2 preview supplies | Still needed |
|---|---|---|
| Performance | Success, fold quality, successful-attempt speed | Failure duration and full workflow throughput |
| Scope | Garments, sizes, scenes, positions, initial configurations | Home count and household-level result distribution |
| Adaptation cost | Fixed checkpoint; zero per-home task adaptation claimed | Operational setup, calibration, and maintenance burden |
| Independence | Public videos and documented company protocol | External replication or third-party audit |
| Intervention | Autonomous retries and recovery included in timing | Remote takeover, manual reset, and abort counts |
| Duration | 785 task attempts | Operating hours, continuous sessions, and multi-day reliability |
| Safety | Disturbance examples are shown | Safety stops, incidents, supervision rules, and qualification evidence |
What the result does not prove
ACT-2’s laundry evidence should not be silently transferred to every behavior shown in the preview. Sunday says the same base model is learning vacuuming, toy organization, zipper fastening, turning pants inside out, and coffee preparation, but explicitly says those capabilities have not yet been tested against the same standard. They remain demonstrations or development examples, not equivalent 99.1% results.
The laundry evaluation also does not establish:
- performance on socks, underwear, accessories, hanging, pairing, or sorting;
- general household manipulation beyond the declared folding task;
- unattended operation around children or pets;
- safe behavior across all disturbances and failure states;
- low remote-assistance or manual-reset rates;
- reliability across software and hardware updates;
- battery uptime, maintenance frequency, or service burden;
- economic value relative to a person or purpose-built appliance;
- customer availability or recurring deployment.
A model can be highly reliable within one valuable capability without being generally reliable across home life. Sunday itself makes a version of this point: a smaller number of dependable tasks may be more useful than many tasks at low reliability.
The fall 2026 beta is the next evidence gate
Sunday says it plans to place Memo with families through a beta program in fall 2026 (beta program). That is a scheduled pilot. Until participating households have operated the robot and results are reported, it should not be described as a completed deployment.
A strong beta report would publish:
- the number of homes, robots, households, and completed operating days;
- scheduled time, powered-on time, task time, and productive time;
- every attempted task, including abandoned, excluded, and safety-stopped attempts;
- remote interventions, manual resets, recoveries, and unresolved failures;
- household-level distributions rather than only one pooled success rate;
- safety events, near misses, contact events, and supervision rules;
- model, robot, and software versions, including updates during the beta;
- task duration and quality under ordinary household use;
- setup, connectivity, charging, maintenance, and support requirements;
- whether families continue using the capability after the novelty period.
These fields would connect a controlled company evaluation to real product reliability without requiring Sunday to disclose proprietary model weights or household identities.
Verdict
ACT-2’s laundry report deserves attention because it includes the information most robot announcements omit: a denominator, declared scope, fixed checkpoint, adaptation cost, category breakdown, grading protocol, quality distribution, speed, and videos.
The evidence supports three conclusions:
- The reported laundry capability is more than a cherry-picked demo. Sunday documents 785 autonomous attempts and seven failures across a broad but bounded garment and scene distribution.
- The zero per-home adaptation claim is technically significant if reproduced. The reported evaluation uses one checkpoint without target-home or target-garment post-training.
- The evidence remains first-party and task-specific. It does not yet prove independent reproducibility, safe sustained home operation, low human support, or general-purpose reliability.
The most defensible summary is: Sunday reports a strong, unusually transparent company-run result for one household manipulation capability. The next test is not another montage—it is a field report from the family beta with complete intervention, safety, duration, and household-level data.
For the broader model landscape, read Robot Foundation Models and Embodied AI. Our guide to reading robot benchmark success rates shows how this 785-attempt result differs from simulated suites, small physical trials, and learning-curve comparisons. For evidence labels that separate demonstrations, pilots, access, and deployment, use our humanoid robot availability tracker. Definitions for autonomy, foundation model, intervention, pilot, and teleoperation are in the Physical AI glossary.
Frequently asked questions
What is Sunday Robotics ACT-2?
ACT-2 is Sunday Robotics’ robot-learning model for its Memo home robot. Sunday presents laundry folding as its first fully evaluated capability and says the same base model is being developed for other household tasks that have not yet received the same evaluation.
Did ACT-2 really achieve 99.1% laundry-folding success?
Sunday reports 778 successful autonomous folds in 785 attempts, or 99.1% with a reported standard error of 0.3%. The denominator, category breakdown, grading process, scope, and videos make the claim unusually inspectable, but it remains a first-party evaluation without independent replication.
Was ACT-2 adapted to each test home?
Sunday says no data from the evaluation homes or their physical garments was used for task-specific post-training or model selection. The same checkpoint and system configuration were used, with no per-home demonstrations or fine-tuning.
Does the ACT-2 result prove a general-purpose home robot is ready?
No. It provides strong first-party evidence for one bounded laundry-folding capability under a declared scope. It does not establish safety qualification, reliability across every household task, unsupervised multi-hour operation, low remote-intervention rates, commercial availability, or recurring home deployment.
How fast was ACT-2?
For the 778 successful attempts, Sunday reports a median completion time of 2 minutes 13 seconds and a mean of 2 minutes 19 seconds, measured from garment retrieval through placement on the folded stack and including autonomous retries and recovery.
What should Sunday report from the home beta?
The most useful field evidence would include every task attempt, excluded or aborted runs, operating hours per home, safety stops, remote interventions, manual resets, recovery outcomes, task speed, model and hardware versions, household-level variation, and retention or continued-use evidence.