
I've been working on some local LLM projects, I pulled my old mining rig back to life and got it working again. The goal was to make a local orchestration agent, that would run different LLMs on different GPUs, so that you could swap them out/upgrade them as needed and get AI capability offline without relying on cloud providers like Claude or Codex. I also added a stack of Orange Pis a friend gave me to act as extra CPU workers. The github is here: github.com/bigbertharig/llm_orchestration
The whole project is a stack of various hardware, scripts, and workers, that you can make plans for from a clear template. The whole thing is run on the idea of "what is the dumbest entity that can perform this task?" I used the cloud LLMs to make the plans and the scripts, one main GPU (3090) that hosts a local brain model, then all the other GPUs either load smaller LLMs, split LLMs, or run scripts, and CPU scripts are passed to the Pis.
For example, one task flow I made was a web scraper. You submit a prompt to the brain, such as "Find me 20 experts in solar energy" or "What happened in the last X days about Y topic?" or "How did people feel about the winter storm in east coast US in Jan 2026". There are different scripts for finding news, finding people, getting information about people (public professional info, like companies worked at, publications, organization participation...), but they all follow the same overall logic:
- Brain reads prompt, determines searches to make
- CPU workers make searches, return lists of results
- Worker LLMs filter results, pick best options
- CPU workers scrape the best options
- Worker LLMs read the results and make a summary
- Brain LLM reads all the summaries and makes final decisions for final report
By cycling all the tasks around and making them all parallel, very large searches can be made quite quickly.
Another task I made was a code analyzer, you can input a Github repo along with some claims of what it does (or it will automatically read the README and other docs), and then the rig investigates all the code to see how well the actual functionality matches the claims.
The exciting thing about this is I've made it very self deterministic. On my other projects, I've been more controlled about the code the AI can write, but with this I made the whole thing in a protected environment. Fresh accounts, fresh machine, its own Github account and gave it full access. There are a few security guardrails I've put in, but mostly it has free reign to change itself and make new projects. This project was also inspired by OpenClaw, I wanted to make a more local version since I didn't trust the unrestrained internet version, so mine has firewall protections and very limited (outgoing only) internet access.
I like the name Prometheus, because this feels like stealing fire from the Gods. Taking the power of the AI overlords and bringing it back home, becuase there may be cheap deals for now to get AI coding access, but as the industry matures and monopolies are locked in, I don't expect those prices and free access to stay!
A problem I'm having now is making an upgrade to how the brain handles the workers. Before I had a clear split between brain LLM and worker LLM, and that worked well, I had 2 models (Qwen2.5-7B and 32B), and everything worked smoothly. I started to add in split GPU capabilities, so that some more advanced tasks could use a smarter LLM, so I added the 14B model and loaded it on 2 GPUs. That became a bit unstable, but it worked.
I then wanted to do more benchmarking, to actually find more models that are better suited for different tasks, but getting all those to load/unload became really confusing. I'm working on porting from Ollama to llama.cpp to get better control over the model management, and then trying to get a library of models so that the brain knows which to pick for each task. I've been fighting it for a few days now, but I'm getting close!