Nubank
Nubank handles varied support requests for card delivery, debt, credit limits, card management, and product explanations. A knowledge-base-only bot cannot reliably use live business data, perform controlled actions, or follow conditional workflows. Offline answer scores alone cannot establish whether real customer satisfaction and self-service outcomes improve.
The team decomposed human support SOPs into independently versioned instructions, routines, tool descriptions, and working memory. The card-delivery agent retrieves profile and logistics data, investigates by address type, and can reissue a card or hand off with context. Human labels, LLM judges, offline simulation, and GEPA judge-prompt optimization guide changes before variants enter small-traffic A/B tests measuring transactional NPS and self-service rate.
Human agents take over low-confidence, frustrated, or unresolved cases with context preserved. Domain experts label evaluation examples and calibrate automated judges. Credit and debt actions remain subject to existing regulated processes, human oversight, and avenues for review. Support data is minimized, pseudonymized, and access-controlled.
- Choose a frequent, measurable card-delivery workflow and establish old-agent and human baselines
- Break support SOPs into maintainable steps, tool permissions, and decision conditions
- Calibrate automated evaluation with expert annotations rather than trusting model self-grading
- Add logistics tools, reissue actions, frustration detection, and concise-response constraints in response to failures
- Start online experiments at 1% traffic and measure both satisfaction and self-service
- Set production boundaries for privacy, contestation, and human takeover
- Offline quality metrics should be tested against online business outcomes
- Satisfaction and self-service can trade off; neither metric is sufficient alone
- Validate the framework in one frequent domain before extending it across support areas
- Tool permissions, handoff, and data protection are as central as prompting
The paper reports 5 production agents. Against previous variants, the card-delivery agent improved AI transactional NPS by 37 percentage points and self-service rate by 29 percentage points, while remaining 10 points below expert human agents on NPS. Debt management remained 23.6 points below expert humans; the product-explainer agent had a temporary 1.5-point self-service decline. No complete ROI or independent audit is published.
- The 100M+ figure is Nubank's customer base, not the number of agent users.
- The 37 and 29 gains are percentage points, not relative percentages; the comparison is against previous agent variants.
- All five improved transactional NPS, but the product-explainer self-service rate temporarily declined.
- Full ROI is not reported; online results and privacy controls remain first-party statements.