Technical Communications · Solutions Architect Interview

One Interface, Many Models

De-risking a production AI assistant in a market that changes monthly

Yariv Barsheshat · September 10, 2026
Case: LLM integration layer for a seller-assistant product — mtl.ai, Montréal

yariv@barsheshat.com1 / 5
The business problem

The product depended on a market
that wouldn't hold still

The situation

  • A production AI assistant for online marketplace sellers — LLM calls at the core of every feature
  • The LLM market shifted monthly: models deprecated within a year, prices varying ~10× across providers and tiers, the "best" model changing hands repeatedly
  • Every provider switch was an engineering project, not a decision — so we were effectively locked in on cost, capability, and availability

Business needs

  • Ship features without re-engineering when models change
  • Control unit cost as usage scales
  • Stay up through a provider outage
  • Adopt a better model the week it lands

Technical needs

  • One internal API over 3 incompatible provider APIs — auth, streaming, tool calls, errors
  • Per-task model choice + quality evaluation before any swap
  • Visibility: who spends what, per feature
One Interface, Many Models2 / 5
The solution I built

Model choice becomes configuration,
not engineering

Assistant featureslistings helper · message drafts · seller Q&A — application code
↓
Unified LLM interfaceone internal API: chat · streaming · tool calls · structured output
↓
Task routerconfig per task: model choice · fallback chain · cost budget
↓
OpenAI adapter
Google adapter
Anthropic adapter
Application code never touches a provider SDK.
Usage meteringtokens & cost per feature, per model
Eval harnessfixed task set replayed before any model swap
Normalized errorsretries, rate limits, timeouts — one behavior
One Interface, Many Models3 / 5
The decision

Three options, honestly weighed

OptionShip speedFlexibilityBurdenVerdict
Integrate one provider directly Fastest today Locked in — swap = rewrite Lowest Rejected
Adopt an open-source framework Fast start Broad, but fast-churning dependency; features we'd never use Their roadmap, our debugging Rejected
Build a thin in-house layer Weeks to build Exactly the surface we need; swap = config We own the adapters Chosen

What the choice cost us — known limitations

  • Lowest-common-denominator risk: normalizing to shared capabilities means provider-specific features need explicit escape hatches
  • Permanent maintenance: every provider API change is our work, forever
  • One more layer to understand when debugging production issues
One Interface, Many Models4 / 5
What it delivered

Model churn became routine ops

3 → 1
provider APIs behind one internal interface
Hours
to swap a model — config + eval run, no rewrite
Per-task
routing to the cheapest model that passes evals

Business impact

  • Features shipped with no knowledge of which provider ran underneath
  • Cost became visible per feature — spend conversations turned data-driven
  • Provider risk stopped being a roadmap item

The lesson I carry

Build the thinnest abstraction that removes the actual risk — and design the escape hatches on day one, because every abstraction leaks.

One Interface, Many Models · Yariv Barsheshat5 / 5
← → navigate · N notes · F fullscreen
1 / 5