Skip to content
SOCIALGOV
civictech

What data.gov is, and what agencies must publish on it

Data.gov is the federal government's public catalog of datasets, and since 2019 a law has required agencies to publish their data there in machine-readable, open formats.

What data.gov is, and what agencies must publish on it
Behind the catalog: agency datasets published from federal data infrastructure.

Data.gov is the federal government's public catalog of datasets, launched in 2009 and run by the General Services Administration. It is not a giant file server — it is a catalog that indexes hundreds of thousands of datasets from agencies across the government, linking you to each one at its source. Since 2019, the OPEN Government Data Act has made the underlying practice mandatory: federal agencies must publish their data as machine-readable, open formats, using open licenses, with metadata describing each dataset. As of 2025, the catalog lists well over 300,000 datasets.

What is the OPEN Government Data Act?

It is Title II of the Foundations for Evidence-Based Policymaking Act, signed on January 14, 2019. The law states, as government-wide policy, that federal data is public information by default and must be published as machine-readable data — data a computer can process directly, in open formats whose specifications are public. It requires each agency to designate a Chief Data Officer, maintain a data inventory, and publish that inventory to data.gov with standardized metadata. The law applies to data the agency chooses to make public; it does not force the release of restricted data, and privacy, security, and confidentiality laws still control what is publishable at all.

What does a dataset listing actually contain?

Every catalog entry is metadata first. A listing names the dataset, describes it, identifies the publishing agency and contact, states the update frequency, and lists the download resources with their formats — CSV, JSON, XML, shapefiles, API endpoints. The entry also carries a usage license or public-domain marking, so you know what you may legally do with the data. When you click through, you land at the agency's own site, where the actual files live. That two-layer design matters: data.gov's catalog is only as good as agencies' inventories, and a dead link on data.gov usually means the agency moved its file without updating its metadata.

How does data get from an agency to the catalog?

Through a documented pipeline. Each agency maintains an enterprise data inventory, and a public listing derived from it is published in a standardized metadata format — GSA's Project Open Data metadata schema, a JSON structure with fields for title, description, contact, licenses, and update cadence. GSA harvests those agency listings on a schedule and indexes them into the central catalog. Agency Chief Data Officers are accountable for the completeness of their inventory, and OMB reviews agencies' data practices as part of the evidence-act framework. This is why coverage varies by agency: everything in the catalog arrives because some office maintained the metadata behind it.

Related stories: What the federal website design standards require agencies to build · What the Plain Writing Act requires of government documents.

What can you actually do with the data?

Most datasets are downloadable files or live APIs, usable in any spreadsheet, database, or analysis tool, and the default license means you can reuse and republish them, including commercially, with attribution where the license requires it. Practical starting points: look up environmental inspection results, browse health statistics by county, pull weather and climate records, or download federal spending extracts that mirror USAspending. If you want to build software on the data, the catalog marks API resources so you can query rather than download. Data.gov itself also offers topic pages — climate, health, education, public safety — that collect related datasets across agencies.

What are the known limits?

The catalog shows breadth but not uniform quality. Datasets arrive at different update frequencies, some listings lag behind the agency's current data, and completeness varies because each agency self-reports its inventory. Datasets with privacy restrictions appear only in summary or aggregated form, and some data — law enforcement records, certain tax data — cannot appear at all. GAO reviews have periodically flagged agencies' uneven compliance with metadata and inventory requirements. None of that changes the default rule, though: when a federal dataset is public, the law's expectation is that you can find it, download it in a format a machine reads, and reuse it without asking permission.

Where did the catalog come from?

Data.gov launched in May 2009, months into the new administration, with a few dozen datasets as a showpiece of the Open Government Directive issued that year. Its legal footing strengthened in stages: the 2013 open data executive order and OMB policy established machine-readability as the default expectation for new federal data, and the 2019 OPEN Government Data Act wrote the principle into statute and created the accountability structure of Chief Data Officers and inventories. The catalog's own technology is open source, and the metadata schema it depends on is published so agencies, states, and other organizations can produce compatible listings.

Is state and local data on data.gov?

The catalog is federal by design; state and local governments run their own open data portals, and hundreds of them now exist, from large city portals to state transparency sites. A few federal datasets include state-level or locally collected data the federal government aggregated — health statistics, education outcomes, climate records — which is often what people are actually looking for when they want "local" data. If your target is municipal data, start with the city or state portal and use the federal catalog for the national picture; both publish in the open formats the federal law pioneered.

How do you find something specific on it?

Search works best when you go in through structure rather than keywords alone. Filter by agency to see everything a department publishes, or use the category and topic pages to browse a domain such as climate, health, or education. Each result card shows the publishing agency and update frequency before you click, which helps you judge whether the dataset is maintained. When a keyword search returns too much, narrow with the format filter — choosing CSV or an API type screens out documents that are nominally datasets but are really PDFs, which are exactly what the machine-readable rule is supposed to phase out.

Frequently Asked Questions

What is data.gov?
The U.S. federal government's public catalog of datasets, launched in 2009 and run by the General Services Administration. It indexes hundreds of thousands of datasets and links to them at their source agencies.
Does the OPEN Government Data Act force agencies to release all data?
No. It requires agencies to publish public data in machine-readable, open formats with metadata, and to maintain inventories and Chief Data Officers. Privacy, security, and confidentiality laws still determine what can be released.
Can I reuse federal open data commercially?
Generally yes. Federal datasets on data.gov carry open licenses or public-domain status, and reuse including commercial use is permitted, following the license terms listed on each dataset.
Who is responsible for the quality of a dataset?
The publishing agency. Each entry names an agency and contact point, and each agency's Chief Data Officer is accountable for the completeness of its data inventory.