AI infrastructure · 3 min read
Pooling GPUs for local LLMs: lessons from a university lab
How we pooled an RTX 5080, RTX 4080, A100 and a 128 GB GPU into one shared cluster for Llama and other local LLMs at NUST, and what to plan for.
At NUST SEECS's Machine Vision & Intelligent Systems Lab, the hardware was there but it was scattered: an NVIDIA RTX 5080, an RTX 4080, an A100 and a 128 GB GPU system, each sitting in its own machine. My job was to make them one shared resource, so that the lab's AI engineers could run Llama and other open models locally for their research. This is what that involved and what I'd tell anyone planning the same.
Why run models locally at all?
- Privacy: research data and unpublished work never leave the lab.
- Cost control: heavy experimentation doesn't generate per-token bills.
- Reproducibility: the exact model version stays put, so results can be repeated.
- Freedom to fine-tune: open-weight models like Llama can be adapted to the lab's own tasks.
The real problem: mixed hardware
Consumer cards like the RTX 5080 and RTX 4080 each carry 16 GB of memory. A data-centre A100 and a 128 GB system can hold far larger models. Treating them as identical wastes the big cards and overloads the small ones. So the first design decision is which models go where:
- Smaller and quantised models on the 16 GB cards, for fast everyday tasks.
- Large models on the high-memory GPUs, where they fit without heavy compromises.
- Splitting a model across GPUs only when it truly doesn't fit, because cross-device traffic costs speed.
One endpoint for everyone
The goal that mattered most to the engineers was simple: one place to send requests, instead of logging into individual machines. A shared gateway in front of the model servers means users don't need to know which GPU runs which model. Most modern LLM serving tools expose an OpenAI-compatible API, so existing code and libraries work with a one-line change.
llama-extract-large), not by the machine they run on. You can then move them between GPUs without anyone's code breaking.Fitting bigger models in less memory
Quantisation (running a model at lower numerical precision) is what makes 16 GB cards genuinely useful. Many tasks such as extraction, classification and cleaning lose very little quality at lower precision, and it frees the high-memory GPUs for the work that really needs them.
What the cluster is used for
- Research experiments: engineers test, compare and fine-tune models for their projects.
- Data extraction: turning messy text, tables and documents into structured data.
- Smarter scraping: models that read each website's layout (see LLM-powered scraping).
- PDF mining: pulling emails, author names and titles out of academic papers.
Lessons for your own setup
- Inventory memory, not just GPUs. VRAM decides what you can run.
- Put a single endpoint in front so users never think about hardware.
- Watch utilisation. The busiest model should live on the best-suited card.
- Pin model versions so research results stay reproducible.
- Keep it inside your network and treat it like any other server: SSH keys, updates, monitoring.
Planning to run LLMs on your own hardware? Let's talk, or see my AI & data automation service.


