McKinsey & Company
Nearly a century of McKinsey knowledge was distributed across more than 40 internal sources, making discovery and synthesis time-consuming. The first platform version also depended on one model provider; as usage grew, cost, speed, and accuracy could no longer be optimized together.
A four-person team built a proof of concept in one week, an MVP in roughly five weeks, and tested it with 200 alpha users before a three-month gradual rollout. Lilli became an orchestration layer combining large and small models, internal and external knowledge, and access controls. The team later rearchitected it as a composable, LLM-agnostic open-source platform and gates new capabilities through alpha/beta testing, user feedback, and quality metrics.
Consultants use Lilli for retrieval, synthesis, and drafting but decide whether output fits the client context and remain accountable for delivery. The product team observes users, prioritizes use cases, runs alpha/beta tests, and manages risk; knowledge contributors improve the underlying corpus.
- Use a one-week proof of concept to validate executive sponsorship
- Prioritize use cases through workshops, interviews, value, and feasibility
- Test the MVP with 200 alpha users before broad release
- Use real unit economics to move from one provider to a multi-model architecture
- Gate capabilities through alpha/beta users and explicit quality metrics
- The product roadmap should start from user problems, not model features
- Single-provider architecture can become a cost and performance bottleneck at scale
- Gradual rollout brings real usage evidence into architecture decisions
- Professional judgment and client accountability remain with consultants
McKinsey reports 72% active adoption, more than 500,000 prompts per month, and up to 30% time savings in knowledge search and synthesis. The team grew from four people to more than 150.
- The 72% adoption, 500,000+ prompts, and “up to 30%” time savings are first-party metrics and do not represent uniform outcomes for every employee or task.
- McKinsey explicitly describes Lilli as a multi-model orchestration layer, not a single RAG instance.