Case study · NUST SEECS · AI infrastructure
One cluster.
Many models.
At NUST SEECS's Machine Vision & Intelligent Systems Lab, I combined the lab's GPUs into one shared compute pool and ran Llama and other open models on it. The lab's AI engineers can use powerful local models for their research, with no cloud bills and no data leaving the lab. I also built LLM-powered scrapers that adapt to each website.
- Lab
- MachVIS · NUST SEECS
- Role
- Full-stack developer & research assistant
- Hardware
- RTX 5080 · RTX 4080 · A100 · 128 GB GPU
- Models
- Llama and other open-weight LLMs
RTX 5080, RTX 4080, an A100 and a 128 GB GPU, combined into a single shared compute pool.
Enough room for large models that won't fit on consumer cards.
Models run in the lab, so research data never leaves it.
One place for the whole lab to run its models.
Weather history scraped from across Pakistan for S2Cool.
Custom scripts for ResearchInn, Scholar Publishing and more, each tuned to its sources.
01 · The cluster
Scattered GPUs, one supercomputer.
The lab had powerful GPUs spread across machines: an NVIDIA RTX 5080, an RTX 4080, an A100 and a 128 GB GPU system. Individually, none could serve the biggest models to a whole team. I connected them into one pool and deployed Llama and other open models on top.
Now the lab's AI engineers point their work at local models backed by the combined hardware. They can experiment, fine-tune and test for their research without cloud costs, rate limits or sending data outside the university.
- Multi-GPU pooling
- Llama
- Model serving
- Fine-tuning
- Testing
02 · LLM-powered scraping
Scrapers that
read like people.
- 01Publisher sitesJournals and publishers, each built differently
- 02LLM reads the layoutThe model works out each site's structure
- 03Per-site scriptsCustom scripts, including for protected sites
- 04PDF extractionEmails, author names, titles and more
- 05Clean datasetStructured by local LLMs, ready to use
How it works
Classic scrapers break whenever a site looks different. Mine use LLMs to understand each website's style and structure, then run custom scripts written for each publisher. That covers sites behind Cloudflare and other anti-scraping protection.
The richest data often sits inside PDFs, so the pipeline opens them and pulls out emails, author names, paper titles and other useful metadata. Local models then clean and structure everything into datasets for research and for targeted academic outreach.
- Python
- Selenium
- LLM parsing
- PDF & OCR
- Anti-bot handling
More from NUST SEECS
climate, live →






