Data Methodology

Data Sources

All data comes from EPA public programs and federal datasets:

Data Vintage

TRI facility and county pages show the 2023 reporting year as the current snapshot; facility trend pages show the full 2019-2023 history. SDWIS data reflects the most recent EPA release. NPL data reflects the current EPA National Priorities List as of the database build date. Each data type is refreshed when EPA publishes updated files.

Processing Pipeline

  1. TRI bulk data files are downloaded from the EPA TRI program data portal. Each record represents one facility-chemical-year combination.
  2. SDWIS data is downloaded from the EPA ECHO system and processed to extract active violations, health-based concerns, and public water system metadata.
  3. Superfund NPL data is downloaded from EPA CERCLIS and filtered to active NPL sites.
  4. All three datasets are linked to geographic identifiers (state, county) and loaded into a structured SQLite database.
  5. State and facility pages aggregate TRI data by chemical, release medium (air, water, land), and carcinogen/PBT classification.
  6. Violation severity is categorized following EPA's health-based vs. monitoring/reporting distinction.

TRI Reporting Thresholds

Not all chemical releases appear in TRI data. Facilities must report only when they:

  • Have 10 or more full-time employees
  • Fall within one of the covered industry sectors (manufacturing, mining, electric utilities, etc.)
  • Manufacture, process, or otherwise use a listed TRI chemical above the reporting threshold (typically 10,000–25,000 lbs/year)

Many smaller facilities and non-covered sectors are not required to report. Absence from TRI does not mean a facility has zero chemical releases.

Data Vintage and Update Frequency

TRI facility and county pages use the 2023 reporting year, the most recent complete year available; facility trend pages show the full 2019-2023 multi-year history. Facilities submit TRI reports annually by July 1 for the prior calendar year, and the EPA publishes compiled data in the fall. SDWIS data reflects the most recent quarterly release from the EPA ECHO system. NPL data reflects the current EPA National Priorities List as of the database build date. Each dataset is refreshed when the EPA publishes updated files.

Accuracy Commitment

PlainEnviro reproduces EPA data exactly as published without editorial modification or advocacy framing. Release quantities, violation counts, and site statuses are presented as reported by the EPA and its partner agencies. Geographic aggregations (state and county totals) are computed directly from the underlying facility-level records. When data is unavailable for a particular geography or time period, the site displays this clearly rather than interpolating or estimating values.

Limitations

  • TRI data is self-reported by facilities. The EPA conducts periodic audits but does not independently verify all reported quantities. Actual releases may differ from reported values.
  • SDWIS violation data reflects violations identified through state and local monitoring programs. Compliance history depends on the frequency and rigor of inspection programs, which vary by jurisdiction.
  • Superfund site status changes as remediation progresses. Current NPL status reflects the most recent EPA update, not real-time environmental conditions at the site.
  • Environmental conditions are complex and cannot be fully captured by any single dataset. TRI, SDWIS, and NPL each cover specific aspects but do not represent a complete picture of environmental quality.
  • Facility counts, release totals, and water-system counts are heavily influenced by how many EPA-regulated entities happen to be located in an area and how consistently they report, not solely by underlying environmental risk. A county with more TRI facilities is not necessarily more hazardous than one with fewer; it may simply host more of the industrial activity TRI covers. Compare figures within the same category (facility-to-facility, county-to-county) rather than treating raw counts as a danger ranking.
  • The county-level environmental burden score is a PlainEnviro-computed composite (facility density, release volume, Superfund concentration, and water-system violations, each capped and weighted, calibrated against the national distribution). It is not an official EPA rating, is not adjusted for population or land area, and does not predict health outcomes. A past violation or an elevated score documents regulatory and disclosure history, it does not itself mean current conditions are unsafe.
  • PlainEnviro presents data without advocacy framing. Readers are encouraged to consult local environmental agencies and public health departments for community-specific context.
  • PlainEnviro is not affiliated with the EPA or any government agency.