Skip to content
Leon Achata
All work
Agents

Multi-agent LLM Gateway

Agents that share one gateway to Bedrock, OpenAI and Gemini — with caching, cost tracking and MCP tools.

The problem

When every agent talks to LLM providers on its own, API keys spread everywhere, costs are invisible and switching models means rewriting code.

How it works

A central LLM gateway puts three providers behind one API, with a TTL cache and per-request cost, token and latency metrics. LangGraph agents call tools through an MCP server and are exposed over REST and WebSocket streaming. Every service runs in Docker, ready for Kubernetes.

Pipeline

  1. 01Client
  2. 02Agent (REST / WS)
  3. 03MCP tools
  4. 04LLM gateway
  5. 05Bedrock · OpenAI · Gemini

Highlights

  • Switch providers without touching agent code
  • Response cache that cuts repeated API spend
  • Real-time streaming over WebSocket
  • Health checks on every service; ready for AWS EKS

Stack

LangGraphMCPFastAPIAWS BedrockOpenAIGeminiDockerKubernetes

Want something like this for your company?

I build production versions of these systems on your data and your infrastructure.

Related service: Agents & workflow automation

Book a call