SMTAQIMZSyed Taqi · Build & Grow

Case study · NUST SEECS · AI infrastructure

One cluster.
Many models.

At NUST SEECS's Machine Vision & Intelligent Systems Lab, I combined the lab's GPUs into one shared compute pool and ran Llama and other open models on it. The lab's AI engineers can use powerful local models for their research, with no cloud bills and no data leaving the lab. I also built LLM-powered scrapers that adapt to each website.

Lab
MachVIS · NUST SEECS
Role
Full-stack developer & research assistant
Hardware
RTX 5080 · RTX 4080 · A100 · 128 GB GPU
Models
Llama and other open-weight LLMs
4GPUs
Pooled into one

RTX 5080, RTX 4080, an A100 and a 128 GB GPU, combined into a single shared compute pool.

128GB
Largest GPU memory

Enough room for large models that won't fit on consumer cards.

100%
Local

Models run in the lab, so research data never leaves it.

1
Shared endpoint

One place for the whole lab to run its models.

15yrs
Of scraped data

Weather history scraped from across Pakistan for S2Cool.

3+
Publisher pipelines

Custom scripts for ResearchInn, Scholar Publishing and more, each tuned to its sources.

01 · The cluster

Scattered GPUs, one supercomputer.

The lab had powerful GPUs spread across machines: an NVIDIA RTX 5080, an RTX 4080, an A100 and a 128 GB GPU system. Individually, none could serve the biggest models to a whole team. I connected them into one pool and deployed Llama and other open models on top.

Now the lab's AI engineers point their work at local models backed by the combined hardware. They can experiment, fine-tune and test for their research without cloud costs, rate limits or sending data outside the university.

  • Multi-GPU pooling
  • Llama
  • Model serving
  • Fine-tuning
  • Testing

02 · LLM-powered scraping

Scrapers that
read like people.

  1. 01Publisher sitesJournals and publishers, each built differently
  2. 02LLM reads the layoutThe model works out each site's structure
  3. 03Per-site scriptsCustom scripts, including for protected sites
  4. 04PDF extractionEmails, author names, titles and more
  5. 05Clean datasetStructured by local LLMs, ready to use

How it works

Classic scrapers break whenever a site looks different. Mine use LLMs to understand each website's style and structure, then run custom scripts written for each publisher. That covers sites behind Cloudflare and other anti-scraping protection.

The richest data often sits inside PDFs, so the pipeline opens them and pulls out emails, author names, paper titles and other useful metadata. Local models then clean and structure everything into datasets for research and for targeted academic outreach.

  • Python
  • Selenium
  • LLM parsing
  • PDF & OCR
  • Anti-bot handling
Next case study Pakistan's
climate, live →