What we can reach, and what we only know is out there

Plate II.

The reachability map

Every federal inspector general publishes their work. That is not the same as anyone being able to read it together. This page is a picture of the difference, drawn from the archive's own records rather than from an impression of them.

The population, and how it was asked

The set here is the 74 inspectors general who are members of the Council of the Inspectors General on Integrity and Efficiency, taken from the contact appendix of its own annual report. Asking them came to 83 hosts rather than 74, and the extra 9 are worth naming rather than absorbing: seven are a second host for an office already counted, and two are not an inspector general at all (oversight.gov, which is the aggregator, and the Smithsonian's main site, whose record in our own files carries a warning that it is the museum and not the OIG). Every count on this page that is about offices is taken over the 74. Every count that is about hosts says so. Among the 74 offices, 40 publish on their own domain and 32 on a path inside the parent agency's site. 6 of the 83 hosts are on oversight.gov itself, which is more than this project's own records say: the field that labels a host's shape tags only two of those six. The hostname is the fact. The label is something somebody filled in.

Each host was asked exactly one question, once: may we collect from you? The answer was read from the host's own robots.txt, literally, as written. Nothing else was fetched in that pass, and no answer was inferred from a neighbor.

Fig. 1.

Fig. 1 as a table: 83 hosts asked
The hostHostsWhat it means
Permits, and answers59Collectable today, on the operator's own terms.
Permits, and does not answer3The policy grants it. The edge in front of the site refuses a script.
Answers, and declines11A clear no, given plainly, and honored.
Neither10Nothing answered, so no policy could be read at all.
The circles are equal in size and the sets are not: this figure shows which hosts fall where, not how large each group is. Two questions that look like one. The operator's policy and the operator's front door do not always agree, and a plan built on either alone will be wrong about 14 hosts. The remaining 10 sit outside both circles: nothing answered, so no policy could be read.

The 3 in the left crescent are the interesting ones. Their robots.txt grants collection in plain terms, and then the edge in front of the site refuses anything that is not a browser. The policy says yes and the door says no. Recording that as PERMITTED would plan collection that cannot happen. Recording it as REFUSED would put words in an operator's mouth that they never said. So it gets its own verdict and stays visible.

On the other side, 11 hosts answer perfectly well and decline. 8 of those are member offices. The other 3 are oversight.gov, the Smithsonian's museum site, and one host whose name this project guessed from a pattern and has never had a person confirm. So the honest sentence is that 8 of the 74 offices decline. A figure drawn over hosts would have said 11, and counted the Smithsonian twice, once here and once inside the permitted set, because its museum and its OIG answer differently.

All 11 of the declining hosts name automated agents in their file. 10 name this project's own agent. Those are two different facts, and the narrower one is the interesting one: the remaining host names other crawlers, and this project is caught by the general rule rather than by name. It declined either way, and it is honored either way.

The 83 verdicts in full
One row per host, one robots.txt each
VerdictHosts
PERMITTED35
PERMITTED-SLOW19
REFUSED-AI11
NO-ROBOTS5
BLOCKED-EDGE4
HOST-UNRESOLVED3
PERMITTED-UNREACHABLE3
BLOCKED-EDGE-TLS2
CHALLENGE1

From asking to holding

Permission is the first step and the cheapest one. Everything after it costs time, and the narrowing below is not a story about refusal. It is a story about how little of the permitted ground anyone has walked.

Fig. 2.

Fig. 2 as a table: the same 83 hosts, five steps
StepHostsWhat was done
Asked83One robots.txt request each, and nothing else.
Permitted62The operator's own file grants it, or declares nothing at all.
Enumerated54Sitemaps read, to find out what is published at all.
Indexed10Every document recorded with its title, type and server.
Copied0A document downloaded and hashed. This step has not happened.
Five steps, one scale. Only 11 of that first drop is anybody declining; 10 more never answered. Everything after it is work this project has not done yet.

8 hosts sit inside the permitted set and have never been visited by any pass, and they are not the ones you would expect. They are exactly the hosts nobody could ask properly: the five that publish no robots.txt at all, and the three whose file grants collection while their edge turns a script away. Not one is a host that said yes and was then simply neglected. What happened to those eight is the next section, and it did not happen by accident.

What visiting turns up

Three passes have been run and they are different instruments. A link walk follows a site's own links and counts files by extension. A sitemap read asks a host for its own index of itself. A document index enumerates every published item and records where it came from. The first two tell you whether a host is worth the third.

The walk covered 17 hosts in 1,514 requests and found 6,761 files. That figure is a floor, not a total: 10 of those 17 hosts hit the stop bound with a queue still unread, and 6 returned nothing at all after a single page. Why those 6 came back empty is not recorded, so it is not claimed here. Not every absence has been explained, and an unexplained one is worth more written down than guessed at.

The sitemap pass is the wider instrument: 54 hosts enumerated, 211,075 URLs read out of 452 sitemap files. It is also the clearest example on this page of why a total needs its parts. That pass counted 44,865 URLs that look like documents, and 40,311 of them, 89.8 percent, came from a single host (www.archives.gov). Only 2 of the 54 hosts produced any documents at all, and 20 returned no URLs whatsoever. Quoted flat, that number would say something about federal oversight publishing that is simply not true.

That pass carries the same truncation caveat as the walk, and its own output does not say so: one host stopped at the sitemap bound with 200 files read, so its count is a floor in exactly the way the walk's is. Files that could not be parsed are counted nowhere, so a host can appear among those that returned no URLs when what happened is that nothing it returned could be read.

A file is not a report. Those counts are files by extension. A PDF can be a semiannual report to Congress, or it can be a poster about phishing. Nobody should quote a number from this section as a count of reports published, and this project will not either, which is why it is stated here rather than in a footnote.

Per host: requests spent, files seen
The link walk, 17 hosts. Files counted by extension, so a PDF here may be a report or may be a poster. This is not a count of reports published.
HostRequests Files, not reports Completeness
www.oig.dol.gov1504,780floor, stopped at the bound
www.ftc.gov150839floor, stopped at the bound
www.flra.gov150351floor, stopped at the bound
www.rrb.gov150189floor, stopped at the bound
www.gsaig.gov150186floor, stopped at the bound
oig.eeoc.gov150121floor, stopped at the bound
www.aocoig.gov150103floor, stopped at the bound
www.oig.dot.gov150102floor, stopped at the bound
oig.usaid.gov15058floor, stopped at the bound
www.uscp.gov15029floor, stopped at the bound
www.tigta.gov83walked out
oig.nsf.gov10stopped after one page
oig.tva.gov10stopped after one page
www.amtrakoig.gov10stopped after one page
www.epaoig.gov10stopped after one page
www.oig.lsc.gov10stopped after one page
www.sigpr.gov10stopped after one page

Who got left out, and on whose say so

29 hosts were skipped by that pass before a single sitemap was read, on the strength of the verdict recorded against them. Most of that is exactly right. 11 hosts declined, 8 of them member offices, and a pipeline that collects from a host that said no is not worth defending.

Fig. 5.

Fig. 5 as a table: the 29 hosts skipped before enumeration
Verdict on fileHostsDid anyone refuse?
REFUSED-AI11Yes. Declined in its own file.
NO-ROBOTS5No. Nobody refused.
BLOCKED-EDGE4Unknown. Nothing answered, so nothing could be asked.
HOST-UNRESOLVED3Unknown. Nothing answered, so nothing could be asked.
PERMITTED-UNREACHABLE3No. Nobody refused.
BLOCKED-EDGE-TLS2Unknown. Nothing answered, so nothing could be asked.
CHALLENGE1Unknown. Nothing answered, so nothing could be asked.
The two rows marked "nobody refused", NO-ROBOTS and PERMITTED-UNREACHABLE, are the ones to look at. They were skipped without anybody having said no.

8 of those 29 hosts never refused anything. Five published no robots.txt at all, which declares nothing and therefore disallows nothing, and three published one that grants collection in plain terms while their edge turns a script away. The pass treated silence as a no, and treated an unreachable yes as a no, and skipped all eight.

That is this project's own conservatism, not a finding about those offices, and the difference matters enough to draw. Failing closed is the right default when the cost of guessing wrong falls on somebody else's website. It is still a choice this project made rather than a fact those offices established, and reporting it as though the offices were out would be the same mistake this page opened by describing.

The gap the whole project is about

On 10 hosts, every published document was enumerated: title, URL, type, and the host that actually serves the file. That came to 20,428 documents.

Of those, none are held here. Not a small number. None. What sits in this project's document directory is 40 of its own bookkeeping files, which are the enumeration written down, plus 12 saved web pages under a single host, and not one of those pages is among that host's enumerated documents. No document has been downloaded. No digest has been taken of one.

Until an hour before this build, this page said 52 were held, because the count was measuring every file in that directory rather than every document in it. The tell was one office showing three where the rest showed four, which turned out to be a missing summary file. A number that moves when a bookkeeping file is written is not measuring the thing its label claims.

Fig. 3.

Fig. 3 as a table: 10 indexed hosts
StateDocumentsWhat that gets you
Enumerated, with a URL anyone can open20,428You can go and read it, one office at a time, while the link lasts.
Held here as a copy, with a digest against it0Nothing. This is the part that does not exist yet.
Not a criticism of anyone. Every one of those documents is already public, and already free to read. What does not exist is the copy that will still be there in ten years, and the index that lets you read across offices instead of one at a time.

Something the enumeration turned up that an impression would not: the documents are not all served by the office whose page they appear on.

Who actually serves the file
20,428 indexed documents, by serving host
Served byDocuments
the office18,417
www.oversight.gov1,840
img.exim.gov46
www.justice.gov12
www.ignet.gov11
www.congress.gov8
www.whitehouse.gov8
www.gao.gov8
about.usps.com6
www.gpo.gov5
www.govinfo.gov5
www.opm.gov5
www.osc.gov5
arc.publicdebt.treas.gov4
www.pandemicoversight.gov3
live-cncs.oversight.gov3
americorps.gov3
www.bop.gov2
osc.gov2
www.ic3.gov2
oig.federalreserve.gov2
www.dea.gov2
www.whistleblowers.org1
www.dhs.gov1
www.brookings.edu1
www.moran.senate.gov1
www.intelligence.gov1
www.cbo.gov1
assets.science.nasa.gov1
www.ltcfeds.com1
home.treasury.gov1
edit.oig.treasury.gov1
www.uscis.gov1
www.fsis.usda.gov1
www.irs.gov1
edworkforce.house.gov1
www.oba.com1
www.us-cert.gov1
www.treasury.gov1
www.cftc.gov1
www.hudoig.gov1
www.fhfaoig.gov1
www.ncua.gov1
www.sec.gov1
nvlpubs.nist.gov1
www.hudexchange.info1
www.gsaig.gov1
oig.justice.gov1
single-market-economy.ec.europa.eu1
www.hsgac.senate.gov1
www.va.gov1
Per host: enumerated against held
The 10 indexed hosts. What this project holds of each host's published documents. The zero is this project's, not the office's.
HostEnumerated Documents held
www.vaoig.gov4,4180
oig.justice.gov3,8930
www.uspsoig.gov3,7290
www.oig.dhs.gov3,0950
www.doioig.gov2,3460
oig.nasa.gov1,0180
www.americorpsoig.gov6390
oig.treasury.gov5330
www.fdicoig.gov4860
eximoig.oversight.gov2710

The refusals, re-asked

On 2026-09-08 this project re-asked every host it had recorded a refusal for, as itself, with the robots file read first and kept. 10 records across 5 hosts were in scope. 5 of those 10 records were a refusal. The rest were not.

Fig. 6.

Fig. 6: 10 recorded refusals, re-asked
HostRecordsWhat it actually was
www.aoc.gov1Nothing was refused. The host answered 200.
www.eac.gov1Nothing was refused. The host answered 200.
www.hudoig.gov2A real refusal. The host answered 403.
www.oig.dot.gov3A real refusal. The host answered 404.
vaoig.gov3Not the host at all. We had built the URL with a space in it and recorded the failure as their refusal.
Of the 5 hosts, 2 refused and 2 never did. The fifth was this project asking a broken question and writing down the answer as though it had been given one.

The last row is the one worth sitting with. Three records said a federal office had disallowed us. What actually happened is that we constructed its sitemap address with a space in the middle, asked for a URL that cannot exist, and filed the resulting failure under the office's name. Nobody at that office did anything. The record was about us and it was written as though it were about them.

None of the original records were edited. The reconciliation supersedes them and sits beside them, which is the only version of this that can be checked rather than believed. A corpus that cannot record having been wrong is not a corpus, it is an assertion.

What each answer is standing on

The last picture is the one this project cares about most, because it is the one that decides whether any of the others can be cited. A verdict is only worth what it can be checked against.

Fig. 4.

Fig. 4 as a table: what the 83 verdicts rest on
FootingVerdictsCan it be re-checked?
A stored digest of the file it was read from65Yes. Re-fetch and compare.
The absence of a file, which is not a refusal5Yes, though absence is all there is to check.
A person opening a browser4Not by machine. No digest, no provenance record.
Nothing. The host never answered9No. There is no answer to check.
Four kinds of footing under one number. They are not equally good, and the honest thing is to draw them apart rather than average them.

4 verdicts rest on a person opening a browser (www.fcc.gov, www.stateoig.gov, www.usda.gov, www.usitc.gov). They carry no digest and no provenance record, because the automated checker cannot reach those hosts at all. They are better founded than what the checker would have written, and they are less checkable than everything around them. Both of those are true, so both are recorded.

That is the thing nobody has followed up on, and it is written down here so that the following up can happen to it.