Data & automation · 3 min read
LLM-powered scraping: why smart scrapers beat brittle selectors
Scrapers break when sites change. How I use LLMs to read each layout, script each publisher and pull emails, authors and titles from PDFs.
For academic publishers and research projects I build pipelines that collect data from dozens of very different websites: journal pages, author listings, archives and PDFs. Classic scrapers rely on fixed CSS selectors and break the moment a site's layout changes. Mine use large language models to understand each page, which makes them far more resilient. Here's the approach.
Why traditional scrapers break
A selector-based scraper says "the author name is in the third div with class meta". That works until the site redesigns, uses a different template for older issues, or loads content with JavaScript. Across many publishers you end up maintaining a fragile script for every layout variation.
Let the model read the page
Instead of hard-coding where every field lives, the pipeline hands the relevant part of the page to an LLM and asks for structured output. The request is along the lines of "return the title, authors, affiliations and contact email as JSON". The model handles layout differences the way a person would. Running the models locally (see our GPU cluster) keeps it fast, private and affordable at volume.
Custom scripts where they matter
LLMs don't replace engineering. Each publisher still gets its own script for navigation: pagination, issue archives, search results, login-free access paths and JavaScript-rendered pages (handled with browser automation such as Selenium). Some sites sit behind Cloudflare and other anti-bot layers, and those need careful, well-paced handling. The script gets to the content, and the model makes sense of it.
PDFs: where the best data lives
In academic publishing the richest details are often inside the paper itself. The pipeline downloads PDFs, extracts their text (with OCR for scanned documents), and pulls out:
- Paper titles and publication details
- Author names and affiliations
- Corresponding author emails
- Other metadata a project needs, such as keywords or sections
Clean before you use
Raw extracted data is messy: duplicates, broken encodings and inconsistent names. A final pass with a local model normalises and de-duplicates the records into a clean dataset ready for research or targeted outreach. This step is what turns "we scraped a lot" into "we have data we can trust".
Doing it responsibly
Scraping power comes with obligations. Before building any pipeline:
- Check each site's terms and robots guidance, and collect only what the project genuinely needs.
- Pace requests so you never degrade a site for its real visitors.
- Handle personal data such as emails in line with data-protection law (GDPR where it applies), and use it only for legitimate purposes.
- Prefer official APIs or data exports whenever a source offers one.
The pipeline at a glance
- Publisher sites, each built differently
- An LLM reads the layout and structures the data
- Per-site scripts handle navigation and protected pages
- PDF extraction pulls emails, authors and titles
- Local models clean everything into a dataset
Need data from sources that don't offer an export? Tell me what you need, or see my AI & data automation service.


