Robotics design has run on one assumption for years: the more human a robot acts, the more people trust it. Engineers keep adding nods, eye contact, and hand gestures to build that rapport. A study led by Hasan Ayaz at Drexel University, published in Science Robotics, finds a serious flaw in that logic. When an expressive robot makes a mistake, trust does not dip. It collapses.
Key finding: People trust an animated, gesturing robot more at first, but a single conversational error cuts that robot’s influence over decisions by more than half. A motionless robot delivering the identical error barely loses ground.
| Metric | Expressive Robot (gesturing, nodding) | Stationary Robot (voice only) |
|---|---|---|
| Initial decision influence | 0.25 | Lower baseline |
| Influence after an error | Drops to 0.10 | Modest decline |
| Brain response to error | Spikes in prefrontal regions tied to norm violation | No comparable spike |
| Oxytocin during error | Rises | Unchanged |
| How the error gets read | Social insult | Technical glitch |
The Experiment: Pepper, Split Into Two Personalities
Fifty adult men each spent about two and a half hours working with Pepper, SoftBank’s humanoid robot, on two collaborative tasks: choosing survival gear for a desert island, then ranking a set of paintings. Half the group got an animated Pepper — gestures, eye contact, nods, occasional “uh-huh” murmurs. The other half got a stationary Pepper giving the same verbal responses without moving a joint. Partway through, the robot started interrupting or offering illogical suggestions to simulate a real conversational mistake.
Researchers measured four things at once: prefrontal cortex activity through wearable fNIRS sensors, salivary oxytocin, self-reported trust on standard questionnaires, and — the part most studies skip — whether participants actually changed their final answer based on what the robot said. Behavior, not just opinion.
Brain Scans Show the “Social Insult” Effect
Once the robot started erring, its sway over people’s choices fell from 0.25 to 0.10 on the team’s influence scale. That part held for both groups. What split the two groups apart was what happened inside the skull.
When animated Pepper made a mistake, activity spiked in the left dorsolateral prefrontal cortex, a region that flags uncertainty and rule-breaking, along with medial prefrontal areas involved in reading other minds. Stationary Pepper’s identical mistake triggered none of that. Same words, same wrong answer, completely different brain response depending on whether the robot had been nodding along five minutes earlier.
This maps onto a well-established idea in communication psychology called Expectancy Violation Theory. Nonverbal warmth raises the bar for what a person expects from an interaction. Break that expectation and the reaction is sharper than it would be from something that never set the bar high in the first place. A vending machine that gives wrong change is a bug. A colleague who does the same thing after weeks of friendly small talk feels like a betrayal. Pepper’s expressive mode put itself in the second category without meaning to.
The Oxytocin Twist Nobody Predicted
Oxytocin gets marketed as the “bonding hormone” — it rises between old friends, new parents, romantic partners. Here it rose at the exact moment the expressive robot broke a social rule, while self-reported trust dropped in the same breath. A hormone associated with closeness moved in the opposite direction closeness would predict, and only in the expressive condition.
Ayaz’s team reads this as social vigilance rather than affection. The brain isn’t leaning toward the robot out of warmth. It’s leaning in to scrutinize a partner that just did something it didn’t expect. Fortune’s coverage of the same study put it bluntly: a friendlier robot loses your trust faster when it messes up.
Why This Matters for Robots Shipping in 2026
This research lands while humanoid platforms are moving out of labs and into homes, warehouses, and retail floors. Tesla’s Optimus, Figure’s line of humanoids, 1X’s NEO, and Unitree’s G1 all lean on some version of expressive, conversational behavior as a selling point, whether that’s natural language interaction, responsive gestures, or both. None of them are Pepper, and none were tested here, but the mechanism this study describes — expressiveness raising the cost of failure — doesn’t depend on which robot is doing the nodding.
Anyone comparing platforms by degrees of freedom or running them side by side on a comparison tool is usually weighing expressiveness as a feature. This study suggests weighing it as a liability too, one that only shows up after the robot has already made a mistake in front of someone.
A Social Repair Protocol, Not a Software Patch
Ayaz’s practical conclusion is sequencing: competence before charisma. A robot’s error handling can’t just be a reboot or a repeated prompt once expressive behavior is switched on, because the brain isn’t processing the failure as a technical event. Three things follow from that:
- Acknowledge the mistake in plain language instead of silently correcting it
- Signal intent — a short explanation of what went wrong reads as more trustworthy than a fast fix with no explanation
- Match the recovery’s tone to the same social register that made the robot likable in the first place, rather than dropping into flat diagnostic language
None of this has been tested yet. It’s an inference from what broke trust, not a proven fix.
What’s Still Unknown
The sample was 50 healthy young men, chosen specifically to avoid the added variance that sex differences and hormonal cycling would introduce into oxytocin measurements. That’s a reasonable methodological call, but it means the findings describe a boundary condition, not a universal rule. Whether the same pattern holds for women, mixed groups, other cultures, or robots with different bodies is still open.
The bigger question sits one step ahead: can a robot apologize its way back into someone’s trust the way a person can, or is machine trust simply more brittle than human trust once it breaks? Ayaz’s team is planning that experiment next, according to Drexel’s own research summary.

