Back to all work

Distributed Local LLM Inference

SwarmAI

An open-source system that turns multiple machines running local Ollama models into a coordinated AI compute swarm.

Built with Python, FastAPI, AsyncIO, Ollama, REST APIs, AWS EC2, ngrok.

The problem

Running LLM workloads locally is limited by the compute available on a single machine.

Why I built it

I wanted to see how far local, self-hosted models could go if several ordinary machines worked together instead of one doing everything.

What I built

SwarmAI coordinates multiple worker machines and distributes AI workloads across the available local models through a central coordinator and scheduler.

How it's put together

Decisions I made

Coordinator / worker model
A central coordinator owns scheduling and state, while workers stay simple: they run Ollama and report health. This keeps the system easy to reason about and debug.
Heartbeats + least-busy routing
Workers send regular heartbeats, so the scheduler knows which nodes are alive and how loaded they are, and routes each job to the least-busy worker.
Async from the ground up
FastAPI with asyncio lets the coordinator hold many in-flight requests to slow model calls without blocking.
Retries on worker failure
If a worker fails or drops mid-request, the job is retried on another node instead of failing the client call.

What was hard

Workers on different networks
Joining machines that aren't on the same LAN required internet-accessible nodes, tested with AWS EC2 and ngrok tunnels.
Measuring honestly
Distribution adds overhead, so speedup is not linear. Benchmarking 1 vs 2 nodes made the real gain — and its limits — visible.

What's in it

  • Coordinator / worker architecture
  • Ollama integration
  • FastAPI APIs
  • Async Python
  • Worker heartbeat
  • Least-busy worker routing
  • Retry handling
  • Agent orchestration
  • Internet-accessible swarm nodes
  • CLI for joining workers
  • Benchmarking

Results

Same workload, time to finish1.66× measured speedup
1 node
113.3s
2 nodes
68.3s

Measured in the project's own test environment. Not a universal performance guarantee.

What I learned

  • Distributed systems fundamentals: health checks, routing, failure handling
  • Asynchronous workloads and concurrency in Python
  • Local LLM infrastructure with Ollama

This project helped me explore distributed systems, asynchronous workloads, model routing, worker coordination, and local LLM infrastructure.

Next project

DevSparkAI Social Hub

AI-Powered Social Media Management SaaS