Case study · Publishing · Data & email · 2025 – present
158K authors,
verified.
Scholar Publishing, a UK open-access publisher with six journals, wanted to reach researchers who had recently published in each journal's field. I built the pipeline that read 186,735 papers across 1,129 publisher sites and delivered 158,251 author contacts, every address verified, matched to the right author and ready for Mailchimp and Brevo. I also run the publisher's journal platform, website, email and ads.
- Client
- Scholar Publishing · UK · six journals
- Journals platform
- OJS 3.3 at scholarpublishing.org/journals
- Publisher site
- WordPress + Elementor at scholarpublishing.org/sse
- My role
- Data pipeline, email, platform, website & ads
Hover the browser to scroll the journals platform
Every address checked deliverable, attributed to the author who owns it, in six import-ready journal segments.
Across 1,129 publisher sites, down to the full text: HTML, OJS galleys and PDFs.
Distinct author emails found before cleaning and verification.
Authors from 2,657 journals and 499 publishers.
In the delivered email column, all returned safe by the verifier.
An audit of the 233,441 contacts already in the account found an older import had put full names in Last Name and years in Phone. All fixed with zero errors.
01 · The brief
The right researchers,
for each journal.
For each of its six open-access journals, the publisher wanted researchers who had recently published related work elsewhere: a verified email, the author's name as they publish it, their institution and country, and the paper that connected them to the subject. It had to import straight into Mailchimp and Brevo, and be clean enough for the team to read row by row.
02 · What I built
Crawl, attribute,
clean, verify.
- A polite, rate-limited crawler that follows each article page to its full text and reads author blocks, affiliation footnotes and corresponding-author lines.
- An attribution layer that decides which author owns which address, instead of attaching one name to every address on a multi-author paper. Every paper was re-read for this: 124,282 papers in one eight-hour pass.
- Column-by-column cleaning for names, affiliations, titles, journals, publishers, years and URLs: measure first, apply second, and back up the master file before every write.
- Subject classification into the six journal areas, with an import file and tag convention per area.
- Third-party deliverability verification of the whole list, with risky addresses held back, not deleted, so the client keeps an audit trail.
On a six-author paper, attaching one name to every address is wrong five times out of six. Matching each address to its owner by surname and local-part evidence replaced 8,739 names with the paper's own attribution and confirmed 111,426.
- Python
- pandas
- PDF text extraction
- Crossref REST API
- Mailchimp Marketing API
- Brevo
- xlsxwriter
03 · Held back, never dropped
Every rejected row
has a reason.
| Why a row was held back | Rows |
|---|---|
| The verifier didn't return the address as safe | 86,583 |
| The name is an initial and a surname that several people share | 9,219 |
| The "author name" is a place, a role or a heading | 6,954 |
| The paper confirms a different address for this author | 3,954 |
| The paper has no year | 3,479 |
| The paper title isn't a title | 1,535 |
| The "title" is a publisher or platform, not a paper | 1,135 |
| The "title" is an issue label or page debris | 719 |
04 · Six segments
One list
per journal.
| Journal | Contacts | Share | Loaded into |
|---|---|---|---|
| British Journal of Healthcare and Medical Research (BJHMR) | 51,006 | 32.2% | Mailchimp |
| Advances in Social Sciences Research Journal (ASSRJ) | 39,908 | 25.2% | Mailchimp |
| Transactions on Engineering and Computing Sciences (TECS) | 26,001 | 16.4% | Mailchimp |
| Archives of Business Research (ABR) | 17,557 | 11.1% | Brevo |
| European Journal of Applied Sciences (EJAS) | 16,057 | 10.1% | Mailchimp |
| Discoveries in Agriculture and Food Sciences (DAFS) | 7,722 | 4.9% | Mailchimp |
Two thirds of contacts published their matched paper in the last two years. Mailchimp audiences were loaded through the API in batches of 500, mapped to existing merge fields and tagged per campaign month.
05 · Data quality
Readable
row by row.
| Column | Filled | Notes |
|---|---|---|
| Author email | 100% | No duplicates, none malformed, all returned safe by the verifier |
| Author name | 100% | As attributed by the paper itself |
| Paper title | 100% | Truncated titles recovered from Crossref |
| Journal | 100% | Page furniture replaced via Crossref and site majority |
| Publisher | 100% | Crossref journal lookup; self-published journals named as such |
| Affiliation | 69.0% | Department and institution; cities, postcodes and biographies removed |
| Country | 78.3% | From the affiliation or email domain only, never guessed |
06 · The publisher behind it
Platform, site,
email & ads.
Scholar Publishing's journals run on Open Journal Systems at scholarpublishing.org/journals, which I set up, migrate, theme and manage, including fixes for issue exports and Google Scholar indexing. Its publisher website at scholarpublishing.org/sse runs on WordPress and Elementor. I set up domain email with SPF, DKIM and DMARC, and run its email marketing plus Google, Meta and Bing Ads campaigns.
- OJS 3.3
- WordPress
- Elementor
- Mailchimp
- Brevo
- Google Ads
- Meta Ads
- Bing Ads
the Khanqah →


