Skip to content
0727
InsightsAcademyCasesEventsSign in
中文/EN
Cases/Nubank
← Back to cases
← Previous16 / 127Next →

Nubank

Five production customer-support agents built through evaluation-driven iteration and A/B rollout
First-party enterprise source Digital banking / financial services Internal enterprise AI deployment ⚙ Technical reference Deep case | Key delivery chain is substantially documented Evidence level A Five production deployments with ongoing A/B iteration
A Evidence level
Building Customer Support AI Agents at 100M-User Scale: An Evaluation-Driven Framework
Publisher: Nubank team / KDD 2026 paper · Technical paper authored by company employees · Direct case-level source
Claim origin: Company project-team disclosure · Independent verification: No · Accessed: 2026-09-23
Evidence level measures whether a source can be located and reviewed; it does not mean vendor-reported claims were independently audited.
Deep case | Key delivery chain is substantially documented 12 / 12
Business context2 / 2
Transformation workflow2 / 2
Technical workflow2 / 2
Human roles & governance2 / 2
Measured outcomes2 / 2
Source traceability2 / 2
All six dimensions meet the current completeness threshold.
Nubank treats support agents as evaluation-driven production systems: modular SOPs and tools, human-calibrated judges, offline simulation, 1% initial traffic, online A/B metrics, and reuse across five domains.
Business problem

Nubank handles varied support requests for card delivery, debt, credit limits, card management, and product explanations. A knowledge-base-only bot cannot reliably use live business data, perform controlled actions, or follow conditional workflows. Offline answer scores alone cannot establish whether real customer satisfaction and self-service outcomes improve.

Solution

The team decomposed human support SOPs into independently versioned instructions, routines, tool descriptions, and working memory. The card-delivery agent retrieves profile and logistics data, investigates by address type, and can reissue a card or hand off with context. Human labels, LLM judges, offline simulation, and GEPA judge-prompt optimization guide changes before variants enter small-traffic A/B tests measuring transactional NPS and self-service rate.

Technical architecture & production workflow
Step 01
Translate human support SOPs into composable, semantically versioned instructions, routines, tool descriptions, and working memory
→
Step 02
Use profile, delivery-status, and card-reissue tools for controlled domain actions
→
Step 03
Build scenario evaluation sets with human labels; calibrate LLM judges against human agreement
→
Step 04
Optimize judge prompts with GEPA in DSPy and check evaluation stability across models
→
Step 05
Launch new variants at 1% traffic before scaling online A/B exposure
→
Step 06
Measure transactional NPS, self-service rate, and task completion while handing uncertain cases to people
→
Step 07
Apply data minimization, pseudonymization, role-based audited access, and retention limits
Key technology & infrastructure components
Modular context engineeringDomain SOP routinesCustomer profile and delivery toolsLLM-as-a-JudgeGEPA / DSPyOffline simulation and human labelsOnline A/B and transactional NPSHuman handoff and privacy controls
Human roles & accountability

Human agents take over low-confidence, frustrated, or unresolved cases with context preserved. Domain experts label evaluation examples and calibrate automated judges. Credit and debt actions remain subject to existing regulated processes, human oversight, and avenues for review. Support data is minimized, pseudonymized, and access-controlled.

FDE delivery actions
  • Choose a frequent, measurable card-delivery workflow and establish old-agent and human baselines
  • Break support SOPs into maintainable steps, tool permissions, and decision conditions
  • Calibrate automated evaluation with expert annotations rather than trusting model self-grading
  • Add logistics tools, reissue actions, frustration detection, and concise-response constraints in response to failures
  • Start online experiments at 1% traffic and measure both satisfaction and self-service
  • Set production boundaries for privacy, contestation, and human takeover
Reusable delivery patterns
  • Offline quality metrics should be tested against online business outcomes
  • Satisfaction and self-service can trade off; neither metric is sufficient alone
  • Validate the framework in one frequent domain before extending it across support areas
  • Tool permissions, handoff, and data protection are as central as prompting
Business outcomes & delivery results

The paper reports 5 production agents. Against previous variants, the card-delivery agent improved AI transactional NPS by 37 percentage points and self-service rate by 29 percentage points, while remaining 10 points below expert human agents on NPS. Debt management remained 23.6 points below expert humans; the product-explainer agent had a temporary 1.5-point self-service decline. No complete ROI or independent audit is published.

Evidence boundaries & verification notes
  • The 100M+ figure is Nubank's customer base, not the number of agent users.
  • The 37 and 29 gains are percentage points, not relative percentages; the comparison is against previous agent variants.
  • All five improved transactional NPS, but the product-explainer self-service rate temporarily declined.
  • Full ROI is not reported; online results and privacy controls remain first-party statements.
Primary source: Building Customer Support AI Agents at 100M-User Scale: An Evaluation-Driven Framework ↗ Additional sources: arXiv abstract, author list, and KDD 2026 metadata ↗
Traceable does not mean independently audited
Open primary public source
0727.ai · Trusted agents, built together.Case research: FDE-case-library ↗ · MIT