Doctolib
Doctolib originally had fragmented data, machine-learning, and engineering platforms, while AI product teams built their own evaluations, model access, and agent frameworks. The first AI-native products took several quarters to reach production. Teams spent time assembling infrastructure instead of product logic, and reuse between teams remained low.
Over more than a year, Doctolib unified its Data & AI Platform and sequenced investment into shared evaluation, governed model access, and framework standardization. A shared evaluation and observability system involves PMs and domain experts; a GenAI gateway switches among compliant models with provider fallback. A common agent framework and golden-path template pre-integrate memory, evaluation, observability, model access, and guardrails so teams can focus on domain logic.
Feature teams own production quality and cost, while PMs and domain experts contribute directly to evaluation. The platform team maintains shared gateways, evaluation infrastructure, and standards. Healthcare safety still requires human governance. The author explicitly says product-asset reuse remains limited and a more advanced agent builder is future work.
- Identify evaluation, model access, and deployment work repeated by early product teams
- Build shared evaluation and a model gateway before standardizing the agent framework
- Bring PMs and domain experts directly into quality evaluation
- Package memory, evaluation, observability, and guardrails into a runnable template
- Validate startup speed with one production beta while tracking reuse gaps
- Sequence platform work by what unblocks AI product delivery
- A unified gateway and evaluation system are foundations for multiple products
- A golden path should start from runnable wiring rather than documentation alone
- Do not generalize one beta's delivery speed to the whole product portfolio
The Doctolib platform lead reports hundreds of evaluation experiments per day. Early AI-native products required several quarters to reach production, while one team at an internal hackathon reached a production beta with an agentic product in about 3 weeks. This single beta does not establish an average delivery time, and no patient outcomes, total ROI, or independent audit are reported.
- About 3 weeks refers to one hackathon team's production beta, not an average across all AI products.
- Hundreds per day refers to evaluation experiments, not production products or users.
- The agent builder and stronger product-asset reuse remain future work.
- The article does not disclose the specific framework vendor, patient outcomes, aggregate cost savings, or independent audit.