Distributed Local LLM Inference
SwarmAI
An open-source system that turns multiple machines running local Ollama models into a coordinated AI compute swarm.
Built with Python, FastAPI, AsyncIO, Ollama, REST APIs, AWS EC2, ngrok.
The problem
Running LLM workloads locally is limited by the compute available on a single machine.
Why I built it
I wanted to see how far local, self-hosted models could go if several ordinary machines worked together instead of one doing everything.
What I built
SwarmAI coordinates multiple worker machines and distributes AI workloads across the available local models through a central coordinator and scheduler.
How it's put together
Decisions I made
- Coordinator / worker model
- A central coordinator owns scheduling and state, while workers stay simple: they run Ollama and report health. This keeps the system easy to reason about and debug.
- Heartbeats + least-busy routing
- Workers send regular heartbeats, so the scheduler knows which nodes are alive and how loaded they are, and routes each job to the least-busy worker.
- Async from the ground up
- FastAPI with asyncio lets the coordinator hold many in-flight requests to slow model calls without blocking.
- Retries on worker failure
- If a worker fails or drops mid-request, the job is retried on another node instead of failing the client call.
What was hard
- Workers on different networks
- Joining machines that aren't on the same LAN required internet-accessible nodes, tested with AWS EC2 and ngrok tunnels.
- Measuring honestly
- Distribution adds overhead, so speedup is not linear. Benchmarking 1 vs 2 nodes made the real gain — and its limits — visible.
What's in it
- Coordinator / worker architecture
- Ollama integration
- FastAPI APIs
- Async Python
- Worker heartbeat
- Least-busy worker routing
- Retry handling
- Agent orchestration
- Internet-accessible swarm nodes
- CLI for joining workers
- Benchmarking
Results
Measured in the project's own test environment. Not a universal performance guarantee.
What I learned
- Distributed systems fundamentals: health checks, routing, failure handling
- Asynchronous workloads and concurrency in Python
- Local LLM infrastructure with Ollama
This project helped me explore distributed systems, asynchronous workloads, model routing, worker coordination, and local LLM infrastructure.
Next project
DevSparkAI Social Hub
AI-Powered Social Media Management SaaS