Back to Portfolio
Open Source / Developer ToolingLive

Swarm Agent Kit

An open-source, production-ready multi-agent orchestration framework for Python backends. Built to fill the gap between simple LLM scripts and real production deployments — where async, persistence, observability, and provider-agnostic routing are non-negotiable.

Author & Maintainer
Feb 2026 – Present
Open Source / DevTools
View Project
100+
LLM Providers Supported
2
Orchestration Modes
PyPI
Published & Installable
MIT
Open Source License

By late 2025 the multi-agent space was crowded but shallow. Every framework I evaluated failed on at least one dimension that matters in production. Some locked you into a single provider — which is fine until your OpenAI bill triples or Anthropic ships a better model. Others had no native async support, which makes them unsafe to run inside a FastAPI or NestJS backend serving real users. Persistence was almost always an afterthought — most frameworks assumed you'd run one agent conversation, in memory, and never restart the process.

Observability was worse. You had no idea which agent handed off to which, what state was mutated, or why a specific tool call happened. Developers building actual products were forced to glue these missing pieces together from scratch every time, and the result was usually fragile.

The bet

Swarm Agent Kit is what you get when you treat the missing pieces as first-class concerns instead of features to bolt on later. Two orchestration modes cover the real design space. Native async safe for production concurrency. A bring-your-own-database persistence layer. Shared state that stays lean as sessions grow. LiteLLM sitting underneath as the provider layer, so you can swap models with a one-line change. And an observability dashboard so you can actually see what your agents are doing.

One pip install, and you're production-ready.

The design decision that changed everything

The hardest design call wasn't technical. It was philosophical.

Every existing multi-agent framework had picked a stance. Either agents plan everything themselves — autonomous, unpredictable, creative — or a central planner controls everything — predictable, less powerful, easier to debug. Both stances have real use cases. Forcing users to pick one at framework level felt wrong to me. A production pipeline where deterministic sequencing matters needs the second. A research workflow where dynamic reasoning matters needs the first. Same team, same product, sometimes on the same day.

So I built both modes into the same core: Unsupervised, where agents autonomously hand off tasks based on context, and Supervised, where a central LLM planner enforces a strict pipeline. Same primitives, same state model, same observability. You switch modes with a flag.

The framework's stance is that there isn't one right stance. Users pick per workflow, not per framework.

Where it came from

The origin was 100Minds.ai. I was building a voice-first AI tutor and coaching platform where different agents needed to hand off tasks reliably inside a FastAPI backend. One agent handling scenario framing, another handling feedback generation, another retrieving grounding content. Every framework I tried either didn't support async natively — which meant blocking the event loop and destroying our sub-second latency budget — or forced vendor lock-in, which meant one API price hike could sink the product.

I ended up writing the orchestration layer from scratch as part of the 100Minds codebase. After shipping, I realized the pattern I'd built was more general than the product it lived in. So I extracted it, cleaned up the API surface, added the LiteLLM provider layer to strip out vendor lock-in, and published it as an open-source library.

State management is where multi-agent systems die

Routing is not the hard part of multi-agent systems. State is. When multiple agents share context, three failure modes emerge fast: state balloons and every LLM call costs more; state gets stale and agents make decisions on outdated context; state gets clobbered because two agents wrote to the same key.

I rewrote the state layer twice before I was happy with it. The final design keeps state as a flat mutable dictionary with namespace prefixes per agent. Agents can read broadly — they see everyone's state — but write narrowly, only within their own namespace. That constraint eliminated the write-conflict class of bugs entirely and gave the observability dashboard something clean to visualize.

Persistence as a hook, not a feature

Every framework I evaluated bundled its own persistence store. Some used SQLite, some used a Redis wrapper, some rolled their own JSON files. All of them were wrong for someone. If you're already running PostgreSQL for your app, you don't want a second store for agent state. If you're on serverless and can't run Redis, a Redis-first framework is dead on arrival.

Swarm Agent Kit's persistence is a bring-your-own hook. You pass in a save function and a load function. That's it. The framework calls them at the right points in the agent lifecycle. Use Postgres, Redis, SQLite, DynamoDB, a JSON file, or your own in-memory dict. The framework doesn't care.

swarm-kit studio — the debugger I built for myself

Building the observability dashboard was where I understood how deeply I'd underestimated the debugging problem. Multi-agent systems fail in ways that look like magic if you can't see the handoffs. An agent hands off to the wrong specialist, the specialist makes a plausible-sounding but wrong decision, and the final output looks reasonable until you notice it's confidently incorrect. Without observability you can't even reproduce the bug.

swarm-kit studio started as a debugging tool for me and became one of the most cited reasons developers use the library. It shows every handoff, every tool execution, every state mutation, in real time, in your terminal. When something goes wrong, you can see exactly why.

What I learned

Building a framework that other developers will actually use forces you to think about API design at a different level. Not "does this work" but "will this still make sense six months from now when someone reads it at 2am trying to fix a production bug." Every parameter I added had to justify its existence. Every default had to be defensible.

State management across agents is deceptively hard. The real complexity isn't routing — it's ensuring shared state stays lean, consistent, and doesn't balloon token usage as conversations grow. Namespace prefixes were the single highest-leverage design decision in the whole library.

Documentation is half the product. A library nobody can figure out how to use is worse than one that doesn't exist — it wastes evaluators' time.

I spent almost as much time on the docs site as on the library itself, and it's paid back multiple times over in adoption.

What I owned

  • Designed the dual-mode orchestration engine — Unsupervised and Supervised — sharing the same primitives, state model, and observability so users can switch modes without rewriting agents
  • Built full async/await support that is safe for deployment inside high-concurrency frameworks like FastAPI and doesn't block the event loop under concurrent sessions
  • Implemented global state management with namespace prefixes so agents can read broadly but write narrowly — eliminating write-conflict bugs and keeping token usage lean
  • Created native bring-your-own persistence hooks — users pass their own save/load handlers for Redis, PostgreSQL, SQLite, or any backend, instead of being locked into a bundled store
  • Shipped swarm-kit studio, a real-time CLI observability dashboard visualizing agent handoffs, tool executions, and state mutations — the fastest way to debug why a multi-agent system did what it did
  • Integrated LiteLLM as the provider layer, unlocking 100+ LLM providers behind a single consistent interface — swap OpenAI for Anthropic or a local model with a one-line change
  • Published on PyPI with full documentation, quick-start guides, architecture diagrams, and MIT license — treating docs as a first-class product surface