Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Expired Domains Scraper API

expireddomains-scraper-api

checks license

The listing on expireddomains.net is fourteen columns wide and thirteen of them are abbreviations. BL, DP, ABY, ACR, RDT. An expired domain scraper that pulls the table without knowing what those mean gives you a spreadsheet you cannot filter, which is the same as no data.

So this repo is mostly a decoder. Every definition below is quoted from the site's own column tooltips, not guessed. Built on ScrapingBee's web scraping API and verified live on 2026-09-15.

It is public, and it is 1 credit

Worth stating early because it is not obvious. The deleted domains listing renders server side, with no login and no challenge:

curl -G "https://app.scrapingbee.com/api/v1/" \
  -H "Authorization: Bearer $SCRAPINGBEE_API_KEY" \
  --data-urlencode "url=https://www.expireddomains.net/deleted-domains/" \
  -d mode=auto

98,324 bytes, spb-cost: 1, 25 rows with all fourteen fields populated. No browser, no proxy tier.

The decoder

Straight from the title attribute on each column header:

Column Header tooltip What you do with it
Domain Domain Name The name, with original capitalisation preserved
BL Majestic External Backlinks Raw backlink count. The headline metric
DP SEOkicks Domain Pop, number of backlinks from different domains Referring domains. Far more meaningful than BL
ABY The Birth Year of the Domain using the first found Date from archive.org Age. - means archive.org never saw it
ACR Archive.org Number of Crawl Results How much history exists. Low means thin
Dmoz Status of the Domain in Dmoz.org Legacy directory listing
C N O D DNS Status .com / .net / .org / .de Whether the matching TLD is taken
Reg Number of TLDs the Domain Name is Registered High means someone owns the name broadly
RDT Number of Related Domains in .com/.net/.org/.biz/.info Typo and variant pressure
Dropped When the domain dropped Freshness
Status Status of the Domain (Available or Registered) Whether you can actually buy it

The pair that matters most is BL against DP. A domain with 1,717 backlinks and a DP of 27 has 1,717 links from 27 places, which is usually one site linking repeatedly. DP is the number worth sorting on.

Selectors

Every cell has a semantic class. No build hashes, nothing that rotates:

rules = {"rows": {"selector": "table.base1 tbody tr", "type": "list", "output": {
    "domain":  "td.field_domain a",
    "bl":      {"selector": "a.bllinks", "output": "@title"},
    "dp":      "td.field_domainpop",
    "aby":     "td.field_abirth",
    "acr":     "td.field_aentries",
    "dmoz":    "td.field_dmoz",
    "com":     "td.field_statuscom",
    "net":     "td.field_statusnet",
    "org":     "td.field_statusorg",
    "de":      "td.field_statusde",
    "reg":     "td.field_statustld_registered",
    "rdt":     "td.field_related_cnobi",
    "dropped": "td.field_changes",
    "status":  "td.field_whois",
}}}

Take BL from the anchor title, not the cell text. The backlink cell also contains two outbound link labels, so reading it as text gives you 17 Majestic.com SEOkicks.de. The a.bllinks element carries the clean number in its title attribute, which is why that one field selects an attribute rather than text.

Live output:

{"domain": "campeonatosmc.com.ar", "bl": "1,717", "dp": "27", "aby": "2008",
 "acr": "198", "dmoz": "-", "com": "available", "net": "available",
 "org": "available", "de": "available", "reg": "1", "rdt": "0",
 "dropped": "Today 09:52", "status": "available"}

That row is what a good find looks like: born 2008, 198 archived crawls, 27 referring domains, and every major TLD still free.

Two number formats to normalise

Both appeared in the same 25 row pull:

  • Thousands separators. bl came back as 1,717. Strip the comma before casting.
  • Abbreviated counts. rdt came back as 1.4 K. Expand the K suffix.

Everything else is a plain integer, - for missing, or one of available and registered.

def to_int(text):
    text = (text or "").strip().replace(",", "")
    if text in ("", "-"):
        return None
    if text.endswith("K"):
        return int(float(text[:-1].strip()) * 1000)
    return int(text) if text.isdigit() else None

Filtering for something worth buying

The columns exist so you can be strict. A reasonable first pass:

def worth_a_look(row):
    return (
        row["status"] == "available"
        and to_int(row["dp"]) and to_int(row["dp"]) >= 10   # real link diversity
        and to_int(row["acr"]) and to_int(row["acr"]) >= 50  # genuine history
        and row["aby"] not in ("-", None)                    # archive.org saw it
        and int(row["aby"]) <= 2018                          # old enough to matter
    )

DP over raw BL, archive crawl count over birth year alone, and always check status, because a listed row is not automatically purchasable.

None of this tells you whether the backlinks are any good. The site reports counts, not quality, so a domain can score well here and still be a former spam network. Check the archive.org history of anything you shortlist before you spend money.

Other listings, same shape

The /deleted-domains/ path is one of several. /domains/ covers the full database and /deleted-com-domains/ narrows to .com, and they share the table markup, so the same rule set works across them. Each page is another 1 credit request, and pagination is a start offset in the query string.

Credit cost

Measured from spb-cost response headers:

Call Credits
A listing page via mode=auto 1
Rejected request 0

mode=auto bills only the rung that worked, which here is the plain rung, and nothing if every rung fails. It cannot be combined with render_js, premium_proxy or stealth_proxy, and sending both returns HTTP 400 while billing nothing. Do not add render_js: the table is in the delivered HTML and rendering would cost 5 for the same rows.

At 25 rows per credit, a full sweep of 10,000 listings is 400 credits. ScrapingBee does not cache, so store what you pull and use the Dropped timestamp to skip pages you have already seen.

Plan tiers are on the pricing page.

Scope

The public listing pages. expireddomains.net gates deeper filtering, saved searches and some columns behind a free account, and none of that is in scope: scraping under login credentials is prohibited by ScrapingBee's terms of service. Everything documented here comes off pages served to an anonymous visitor.

Metrics on these pages are third party estimates from Majestic, SEOkicks and archive.org rather than measurements by expireddomains.net, so treat them as a shortlist signal and verify anything you intend to buy. Acquiring a domain for its backlink profile sits in a grey area with search engines, and nothing here is advice about that.

Reference: extraction rules, data extraction feature.

Adjacent endpoints: domain AU API, Google my business scraper API, Bing search API, DuckDuckGo search API.

FAQ

Do I need an account to scrape expireddomains.net? Not for the public listings. All fourteen columns came back on an anonymous request at 1 credit. Deeper filtering behind the account wall is out of scope.

What is the difference between BL and DP? BL is total backlinks, DP is the number of distinct linking domains. Sort on DP. A high BL with a low DP is one site linking many times.

Why is my backlink column full of text? Because the cell holds the number plus two link labels. Read a.bllinks and take its title attribute instead of the cell text.

What does a dash mean? Missing data. On ABY it means archive.org has no record of the domain, which usually means it was never meaningfully published.

Is 1.4 K a real value? Yes, in RDT. Some counts are abbreviated with a K suffix and need expanding before you compare them numerically.

License

MIT. See LICENSE.