THE SIGNAL IN ONE SENTENCE
Meta launched Muse as a personal AI agent that can act on a person's behalf. It can browse, fill forms, send messages, shop and, more recently, call US businesses to book appointments, check inventory or negotiate a bill. Reuters has now reported that Meta tested a second route for some employee calls: Muse could hand the request to a trained human contractor, who placed the call and worked it through. The problem was not the existence of a human fallback. Humans rescue automated systems all day. The problem was that the handoff was quiet. Reuters reviewed internal company posts in which a Meta executive acknowledged that beginning contractor-placed calls without proper disclosure was a miss and said the feature had been rolled back for now. The same posts described tests in which human-made calls reached a claimed 95 to 98 percent success rate, compared with a lower percentage for AI calls. That result is not an autonomy score. It is a blended service score produced by a person working behind an AI interface. The distinction matters to both ends of the telephone. The Muse user needs to know when a contractor can see the request, transcript or account details. The person answering at a salon, insurer or cable company needs to know whether they are speaking with software, a contractor or a contractor speaking for a customer through software. The worker needs clear instructions, access limits, supervision and a way to refuse unsafe or discriminatory tasks. Meta's public launch and security documentation describes an isolated cloud computer, protected credentials, a permission authority called Sentinel and an audit trail of agent actions. Those are serious software controls. They do not explain the contractor test, who could access which data, how the human handoff appeared in the audit trail or what the call recipient was told. Meta says the employee response was overwhelmingly positive, that the test was intended to improve safety and privacy before public release and that any rollout would wait until the feature was ready and included proper disclosures. That is a useful promise. The plain signal is simpler: human assistance can make an AI service more reliable, but the human must not become a hidden benchmark boost, an invisible privacy boundary or a stranger wearing the machine's name. Put the handoff on the screen, say it on the call and count it separately.
01
WHAT ACTUALLY CHANGED
Reuters reported on September 22 that Meta had tested a human concierge for Muse, using contractors to handle some phone calls placed through the personal agent.
The report is based on internal company posts reviewed by Reuters and direct responses from Meta. Meta has not published a complete primary report describing the contractor test.
Muse's calling feature can contact US businesses for errands such as booking a haircut, checking whether an item is in stock, gathering contractor quotes or negotiating a service bill.
Meta began employee testing of phone calls in August and gradually made AI calling available to users after Muse launched on September 8, according to the internal posts reported by Reuters.
The posts said Meta enabled the human concierge route for half of its employees during the test. Employees could join an opt-out group.
A trained contractor could receive a request, place the call and work it through. That is a human-delivered service initiated through an AI product, not an autonomous call completed by the model.
A Meta executive acknowledged internally that testing contractor calls without proper disclosures was a miss and said the feature had been rolled back for now, Reuters reported.
Internal tests put the success rate for calls made by humans at 95 to 98 percent, compared with a lower percentage for AI calls. Meta did not publish the test design, sample size, definition of success, comparison rate or confidence interval.
One employee said a transcript of a call to an internet and cable provider showed the contractor making a racist reference. The Meta executive apologized and said that worker would not work on Meta projects again.
Another employee reported that an insurer repeatedly ended calls after hearing that Muse was an AI. The contractor route was intended in part to overcome resistance from people answering automated calls.
Meta spokesperson Daniel Roberts told Reuters the employee response was overwhelmingly positive and that the test was gathering feedback before any public release.
Roberts said Meta was working with merchants and would release the potential calling feature only when ready and with proper disclosures.
Meta's September 8 launch page says Muse users control connected apps, approve sensitive actions and receive an audit trail. It also says conversations and virtual-machine data are not shared with Meta's ad systems.
Meta's technical safety page describes an isolated Linux environment, separated credentials and Sentinel as the sole authority for network access and connector actions. It does not document the human concierge route.
Sensor Tower data cited by Reuters put Muse above 2.5 million downloads since launch. Adoption does not establish the reliability, privacy or disclosure quality of the calling test.
02
WHY THIS MATTERS
A human fallback is not automatically a scandal. It can be the safest way to recover when an automated caller is confused, a business refuses to engage or a request becomes sensitive. The useful question is whether the handoff is disclosed, permissioned and logged.
Autonomy is a product claim and a measurement category. If a person completes a task, the outcome should appear in a human-assisted column. Combining it with model-only calls makes the machine look more capable than the evidence supports.
Task success is not the only outcome. A call can end with an appointment booked while still violating the user's privacy, misleading the recipient, mistreating a worker or recording a discriminatory statement.
There are two consent moments. The customer can authorize Muse to make a call, but the person answering also deserves to know the identity and nature of the caller. One person's app permission does not settle the other person's disclosure rights.
The contractor adds a new data boundary. A request about an insurance claim, account problem, medical appointment or household bill may expose names, addresses, service details and fragments of personal history. Software isolation does not answer which of those details a human can see.
The public security architecture is still relevant. Keeping credentials away from the model and screening network actions can reduce certain technical risks. But controls built around software components do not automatically govern a call-center workflow.
An audit trail should reveal the route, not merely the final action. A user needs to see when Muse acted alone, when a person took over, what information crossed the boundary and what the person said or changed.
Disclosure must be timely. A sentence buried in terms of service does not help a customer decide whether to send a sensitive request. Nor does it tell a merchant at the start of a live call who is speaking and for whom.
The receiving business has legitimate operational concerns. A salon may accept an automated booking call. An insurer or bank may require identity checks that an AI or contractor should not bypass. A disclosed caller lets the business apply the correct process.
Workers are part of the system, not a disposable patch. Training, pay, monitoring, cultural competence, access controls, incident review and appeal all affect product quality. Removing one contractor after a harmful remark does not substitute for a documented labor system.
The test also exposes a market tension. People want agents to complete annoying calls, while many recipients distrust robocalls. Secretly substituting a person can raise completion, but it sidesteps the trust problem instead of solving it.
The most credible product would offer visible modes. Users could choose AI only, human fallback allowed or human only, with different privacy terms, prices and expected completion rates. The recipient would hear the selected mode before substantive conversation begins.
This is a material follow-up to Muse's launch, not a repeat of the launch story. The new event is evidence that the advertised agent system briefly depended on an undisclosed human route for some calls and that Meta rolled that route back.
03
WHERE IT COULD HELP
- Show a clear human handoff notice before a request leaves the user's device, including which data the contractor can view and whether the call will be recorded or transcribed.
- Begin every call with a short spoken disclosure naming the service, the customer being represented and whether the speaker is an AI, a human contractor or a combination.
- Require affirmative user consent for human access instead of placing employees or customers in a test unless they find an opt-out group.
- Separate model-only, human-assisted and human-only completion rates. Publish the failure definition, sample size, task categories, retries, time, cost and customer correction rate for each route.
- Add the handoff to the audit trail with a timestamp, reason, contractor role, data fields exposed, actions taken and final disposition.
- Use minimum necessary disclosure. A contractor booking a haircut should not receive unrelated email, calendar, payment or account data stored in the agent environment.
- Block contractors from seeing raw passwords, payment credentials and authentication secrets, and prevent them from asking a recipient to weaken identity checks.
- Give users a mode that forbids human fallback for sensitive categories such as health, insurance, finance, legal matters and intimate communications.
- Give call recipients a simple way to refuse automated or contractor-mediated calls and to reach the represented person when identity or authorization matters.
- Train and audit contractors for nondiscrimination, impersonation, privacy, escalation and local call-recording rules. Publish aggregate incident and remediation data.
- Test whether the recipient understood the disclosure rather than measuring only whether a sentence was played. Comprehension is the safety property.
- Let independent researchers inspect a documented version of the handoff workflow, including redacted transcripts, permission boundaries and failure cases.
- Price the service honestly. If reliable completion requires paid human labor, include that cost in the product model instead of treating it as invisible scaffolding for an autonomous claim.
KEEP A HAND ON THE WHEEL
The contractor test is documented through Reuters' review of internal Meta posts and Meta's responses, not through a public test report from the company. Reuters reported that the route reached half of employees during one testing phase and that employees could opt out. This does not mean half of all Muse users received human-made calls, that every call used a contractor or that the feature reached public production. The 95 to 98 percent success range is an internal company figure without a published protocol, sample size, task mix, baseline, error bars or independent audit. The report does not establish the contractors' employer, location, pay, staffing level, exact data access, retention rules or full scripts. Meta's launch and safety pages verify the advertised software architecture and user controls, but neither page documents the concierge test. The reported racist reference is one incident described by an employee, not evidence about every contractor. Meta says the feature was rolled back for now and any future release will include proper disclosures. Watch for a public handoff specification, consent flow, contractor data map, spoken recipient notice, separate performance metrics, labor standards, incident reporting and proof that the audit trail identifies who actually completed each call.
04
TERMS WORTH KEEPING
SOURCES AND VERIFICATION STATUS
This article was written from the materials below. Product claims and dates were checked against those sources on September 23, 2026.
PUBLICATION RECEIPT: Revision 1. Published September 23, 2026.
THE PUBLICATION ENGINE
WANT A SIGNAL OF YOUR OWN?
We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.
WORK WITH US