US AI Law Database

Datasheet for Datasets

Datasheet: US AI Law Database, release v0.6.0-2026-09-13

Prepared September 14, 2026. Maintained by the AI Governance Collective (Kyle David, editor).

About this document

A datasheet is a structured description of a dataset, written by the people who built it, that answers a fixed list of questions: why the data exists, what it contains, how it was gathered and checked, what it should and should not be used for, how it is distributed, and how it will be maintained. The format follows "Datasheets for Datasets" (Gebru et al., 2021). The questions below are the paper's, in the paper's order; the human-subjects questions that do not apply are noted and skipped. Where the project's records do not answer a question, the answer says "Not applicable" or "Not yet determined" rather than guessing.

This datasheet describes release v0.6.0-2026-09-13 of the US AI Law Database, the structured companion to the US State AI Laws library. Every page of the database shows its release tag. When a new release ships, this datasheet will be revised to match.

1. Motivation

For what purpose was the dataset created? Was there a specific task in mind? Was there a specific gap that needed to be filled?

The database exists so that AI governance professionals (GRC, legal, privacy, and AI/ML practitioners) can answer specific questions about United States AI statutes without re-reading every act: which laws reach a developer versus a deployer, what a law requires and by when, who enforces it, whether a private party can sue, and what the penalty unit is. The gap is structural. Trackers and alerts summarize, rarely cite the section, and tend to flatten scoped provisions into absolutes. This database was designed so that every cell traces to a section of the enacted text and carries a flag saying how it was verified.

The underlying library began as course material on state AI law and became a library-first effort. The database is its queryable form, and the project's stated direction is to grow it toward a US-focused retrieval database for an agent.

Who created the dataset and on behalf of which entity?

Kyle David ("Dr. David"), founder of the AI Governance Collective (AIGP, CIPP/US, CIPP/E, CIPM, FIP, CISSP; ISO 42001 lead auditor), on behalf of the Collective. Kyle sets the standing rules, decides which laws and bands are admitted, and reviews every batch. AI assistance (Claude) is used for text extraction, drafting, cross-referencing, and the adversarial fact-check audits, under those rules and with the verbatim statutory text as the controlling source.

Who funded the creation of the dataset?

Not yet determined in any public record. The project documents name no grant or outside funder; the only funding arrangement they describe is that the database is a benefit of Collective membership.

Any other comments?

The database is a reference for governance work, not legal advice. Anyone relying on it for a legal position should verify against the official source, which each record links.

2. Composition

What do the instances that comprise the dataset represent? Are there multiple types of instances?

Five tables hold five instance types. An instrument is one enacted law, or one multi-bill package treated as a single law (Michigan's Public Acts 263 through 266). A duty is one typed obligation a law imposes. A term is one verbatim statutory definition. An event is one dated occurrence for a law. A source is one external document used by any field.

How many instances are there in total?

105 law records, of which 98 count as current in-scope instruments and 7 are excluded but retained for reference: the repealed Colorado SB 24-205, the superseded original New York RAISE Act, the Tennessee study act SB 1700, the Tennessee authentic-image context note SB 2041, and three struck or enjoined election-deepfake statutes (California AB 2839, California AB 2655, Hawaii Act 191). 104 records are state instruments and 1 is federal, the TAKE IT DOWN Act, admitted because it anchors the state cluster of non-consensual intimate imagery laws. The duties table has 422 rows across 13 duty types; the terms table has 259 definitions; instrument records carry 198 chamber roll-call rows and 24 rulemaking delegations. Row counts for the events and sources tables are not reported in the release notes.

The 98 current instruments are organized by regulatory model into ten bands:

BandNameInstruments
1Frontier Safety3
2Consequential Decisions5
3Transparency5
4Interaction and Companions15
5Health Care17
6Synthetic Media and NCII43
7Omnibus1
8Algorithmic Pricing5
9Personhood and Legal Status2
10Audit and Assurance2

Band 10 was founded with this release on September 13, 2026. Band 6 includes a complete 33-note census of election-deepfake laws across 30 states, 2019 through 2026.

Does the dataset contain all possible instances or is it a sample of instances from a larger set?

It is a curated set: not a sample, and not a census of every statute that mentions AI. All fifty states were screened for enactments from 2019 through September 2026. Admission follows standing rules: only signed laws are admitted (pending bills stay on a watch list); struck or enjoined laws are kept as non-counted reference records; federal statutes enter only where they anchor a state cluster. Scope is statutes only, with no regulations, executive orders, or bills. Election-deepfake laws are in. General privacy laws, criminal-code AI amendments, and neurorights laws are out, and those scope decisions are parked rather than settled. Within the scope it claims, the set is treated as complete as of the release date; a reconciliation gate fails the build if the count of current instruments disagrees with the library's index.

What data does each instance consist of?

Each instrument record carries identity (state, bill, chapter, codified citation, official source URL), dates (signed, effective, operative, sunset), status, band, covered roles and the statute's own actor terms, numeric thresholds, enforcer, private right of action, penalties with units, cure periods, rulemaking delegations, litigation posture, sponsor party, governor party, trifecta at enactment, and roll-call votes. Each duty row carries a type, the actor role it binds, its trigger, its deadline or cadence, a scoped description, a citation into the library note, and a confidence flag. Each term row carries the defined term and the statute's exact words. The database does not contain the full text of each act; that lives in the private library notes every record cites.

Is there a label or target associated with each instance?

Not in the machine-learning sense. Several fields are editorial classifications and are flagged as such: the band, the sub-cluster within a band, and the definitional family a law's AI definition belongs to (federal-style, EU-style inference, technique-based, generative-output, or other). Duty types come from a fixed 13-value vocabulary: prohibition, disclosure, protocol, reporting, notice, consent, documentation, audit, takedown, registration, detection-tool, contract-terms, and other.

Is any information missing from individual instances?

Yes, and the schema marks it. A value known to exist but not yet captured is flagged "pending" and compiles to null plus the flag; this release reports zero pending fields. Some absences are real: five laws passed by voice vote and have no floor tallies, and five committee or budget bills have no individual lead sponsor. Penalties that ride an external statutory schedule, such as a state's unfair and deceptive practices act, are described in words rather than given as figures the AI statute does not state. The schema has no per-duty expiry field, so two California SB 243 duty rows that cease to be law on January 1, 2027 remain in the duties table with the expiry recorded only in the record's notes.

Are relationships between individual instances made explicit?

Yes. Duties, terms, and events reference their instrument by a stable identifier. A superseded law points at its replacement (New York RAISE 2026 to RAISE 2025). Litigation entries carry a stable case slug shared with the events table so a future case-law dataset can join on it. Multi-bill instruments list every constituent bill.

Are there recommended data splits?

Not applicable. This is a reference dataset, not a training corpus.

Are there any errors, sources of noise, or redundancies in the dataset?

Known ones are listed here. Several 2026 California texts are single-hosted on LegiScan because the official site blocks automated retrieval, and their chapter numbers come from aggregators pending official posting. Some vote tallies conflict between sources and are recorded with both figures (SB 867 in the Senate, 40-0 versus 39-0). Effective dates for non-urgency California acts are inferred from the constitutional default of January 1 following enactment where the text is silent, and are labeled as inferred. The covered_roles field uses "other" for auditors, verification bodies, and advertisers because the role vocabulary predates those bands. Litigation posture is secondary-sourced and dated.

The error pattern the audits hunt is compression that flattens scoped provisions into absolutes: dropped intent elements, dropped qualifiers such as "in violation of state or federal law," wrong penalty units, and effective versus operative dates conflated. The audit for this release corrected every finding it raised, but the pattern is a standing risk in any hand-authored legal summary and readers should expect residual instances.

Is the dataset self-contained, or does it link to or otherwise rely on external resources?

Every record carries the official legislature or code URL and, for political fields, a link to the legislature or vote record. There is no guarantee those pages stay constant, but the database does not depend on them at query time: the values were transcribed, the private library keeps the verbatim text as it existed at extraction, and each release is a frozen snapshot. The statutory words belong to the public. The mirrors used for some texts (LegiScan, state.public.law, Justia, FindLaw) have their own terms of use, which have not been assessed for downstream reuse.

Does the dataset contain data that might be considered confidential?

No. Everything derives from enacted law, public legislative records, and public court filings.

Does the dataset contain data that, if viewed directly, might be offensive, insulting, threatening, or might otherwise cause anxiety?

Some of it, in statutory language. Band 6 covers non-consensual intimate imagery and sexual deepfakes, and the definitions and duty descriptions quote the statutes' terms for sexual conduct and intimate imagery, including provisions concerning minors. Band 4 companion-chatbot laws describe required crisis protocols for suicidal ideation and self-harm. The text is legal rather than graphic, but readers working in those bands should expect it.

Does the dataset identify any subpopulations?

Not in the demographic sense. Political metadata records the party of a bill's primary sponsor, the party of the governor who signed it, and the state's trifecta at enactment.

Is it possible to identify individuals, either directly or indirectly, from the dataset?

Only public officials and litigants in their public roles: governors from signing dates and states, lead sponsors from the cited legislative records, and litigants from case names (xAI v. Bonta). No private individuals appear.

Does the dataset contain data that might be considered sensitive in any way?

The party affiliation of public officials as recorded in public legislative records. Nothing else in the sensitive categories applies.

Any other comments?

The database is deep rather than exhaustive by design. Any summary or visual derived from it should carry a "laws in this database, as of release date" caveat.

3. Collection Process

How was the data associated with each instance acquired? Was it directly observable, reported by subjects, or inferred or derived from other data? If derived, was it validated?

The data comes from a three-level source hierarchy. First, enacted statutory text, transcribed verbatim from official legislature sites or, where those block automated retrieval, from mirrors, with dual-source cross-checks on every operative passage; statutory words are never edited, and drafting quirks are preserved and noted. Second, the library: 105 Markdown law notes, each with an "At a glance" summary that cites sections, cross-links, and the full verbatim text. The library is the source of truth. Third, the database: one YAML record per law, authored from the notes, with every legal-fact field citing its note section and block anchor. Political fields cite legislature and vote records directly, never the note, and never carry the verified-verbatim flag.

Every field carries one of five confidence flags: verified-verbatim (quoted from the embedded statute text), dual-sourced, single-sourced, editorial (a classification judgment), or pending. Validation is by adversarial audit, described next.

What mechanisms or procedures were used to collect the data? How were they validated?

Text retrieval used web search and fetch only; no scraping scripts. Extraction, note drafting, and record authoring were done by AI agents working from the verbatim texts under written rules, with the notes as the sole source for legal fields and no web access during record authoring. A compile script (compile.py) is the only path from the YAML records to the SQLite, CSV, JSON, Parquet, and Postgres outputs, and derived files are never hand-edited. The compile validates every record against a schema, checks that every cited note file and block anchor exists, and runs the reconciliation gate against the library's index.

The process is validated by adversarial fact-check. Every batch of notes and every release is reviewed by an independent agent instructed to refute each claim against the verbatim text; findings are graded FATAL, MATERIAL, or MINOR and applied by fix scripts that abort if any text anchor fails. The audit for this release (September 13, 2026) returned 47 findings across the batch: 0 fatal, 8 material, 39 minor, all corrected. Earlier database audits found 3 fatal errors in total (an effective date copied from the wrong field, fixed in the extraction script so the fix survives regeneration, and two vote tallies), all corrected. Each audit's log is appended to the library's decision tree and to the dataset's PROVENANCE.md.

If the dataset is a sample from a larger set, what was the sampling strategy?

Not applicable. Admission is deterministic under the standing rules in Section 2.

Who was involved in the data collection process and how were they compensated?

Kyle David, as editor, and AI agents (Claude) under his direction. No students, crowdworkers, or contractors were involved.

Over what timeframe was the data collected? Does this match the creation timeframe of the data associated with the instances?

The laws were enacted between 2019 and September 2026. The database layer was authored between schema approval on July 28, 2026 and this release on September 13, 2026, in six numbered releases, on top of library notes built over the preceding months. The newest records are California's September 9 and 10, 2026 signings (SB 1119, SB 867, SB 813, AB 1405) and New York's synthetic performer advertising law, in force June 9, 2026.

Were any ethical review processes conducted?

Not applicable. The data is public law and public legislative record.

Does the dataset relate to people?

Only incidentally, through public officials and litigants named in public records. The remaining human-subjects questions in this section (direct collection, notification, consent, revocation, impact analysis) are not applicable.

4. Preprocessing, Cleaning, and Labeling

Was any preprocessing, cleaning, or labeling of the data done?

Yes, at two levels. Between the statute and the library note, the enacted text was transcribed verbatim and normalized into Markdown with every delta accounted for by formatting only. Between the note and the database, editors assigned bands, sub-clusters, and definitional families (all flagged editorial); typed each duty against the 13-value vocabulary; normalized actor roles to seven values (developer, deployer, operator, platform, provider, any-person, other) while keeping the statute's own actor terms verbatim alongside; encoded penalties as amount, unit, and per-what with the original prose preserved; and derived the "counted" flag mechanically from status. Audit fixes go into the extraction scripts or the YAML records, never into compiled outputs.

Was the "raw" data saved in addition to the preprocessed data?

Yes. The verbatim text extracts and the library notes are retained privately as the source of truth, and each record names its note file and, where one exists, its extract file. The raw data is not distributed with the database; the official source URL on each record points to the public statute.

Is the software that was used to preprocess, clean, or label the data available?

Not yet determined. The compile and validation scripts exist and are described in the project's specification, but no decision has been recorded on publishing them.

5. Uses

Has the dataset been used for any tasks already?

Yes. It powers the members' web application at laws.aigovcollective.com, where members browse by band, state, status, enforcer, and private right of action; open a full record per law; search the 422 duties in plain language; look up the 259 definitions; run 15 ready-made views (from a census by band and a status board to penalties, litigation watch, and trifecta at enactment); and download any view as CSV. Internally it also runs in a Datasette workbench, and a compiled covered-actors view is regenerated from it each release.

Is there a repository that links to any or all papers or systems that use the dataset?

No.

What (other) tasks could the dataset be used for?

Applicability screening (which laws reach an organization given its role, thresholds, and states); comparative analysis of how statutes define AI, allocate enforcement, or structure penalties; tracking effective and operative dates; and retrieval-augmented question answering by an agent, for which the compile emits Parquet files.

Is there anything about the composition of the dataset or the way it was collected that might impact future uses?

Several things. It is a reference, not legal advice, and legal use should verify against the official source. Scope is statutes only, and named categories are out, so a "no law found" result inside the database is not evidence that no law applies. Editorial fields are classifications, not statutory facts. Confidence flags should travel with any downstream use: a single-sourced field deserves less weight than a verified-verbatim one. Effective is not operative; the two dates are separate fields for a reason. The duties table should be read with the release notes, since the schema cannot yet express per-duty expiry.

Are there tasks for which the dataset should not be used?

It should not be the sole basis for a legal opinion, a compliance certification, or a filing. It should not be used to train a model to generate statutory text, since it holds definitions and scoped descriptions rather than full acts and would teach a model the compression the project works to avoid. Political metadata should not be used to characterize individual legislators beyond what the cited public record states.

6. Distribution

Will the dataset be distributed to third parties outside of the entity on behalf of which it was created?

Yes, to members of the AI Governance Collective. Access is included with membership at the Foundation tier, which every member holds. Non-members see only the public landing page, which shows the census counts and the release tag.

How will the dataset be distributed? Does it have a digital object identifier?

Through the web application at https://laws.aigovcollective.com, with sign-in by email link using the address on the member's Collective account (no password), and through per-view CSV downloads from that application. Releases are identified by a numbered, dated tag (v0.6.0-2026-09-13), not by a DOI; no DOI is recorded.

When will the dataset be distributed?

It is available now, at the release this datasheet describes.

Will the dataset be distributed under a copyright or other intellectual property license, or under terms of use?

The project's stated terms: the verbatim statutory text belongs to the public; the compilation, editorial layers, and structure are the library's work product; the database is an educational reference, not legal advice; verify against official sources for legal use. A formal license or terms-of-use document for members is not yet determined.

Have any third parties imposed IP-based or other restrictions on the data associated with the instances?

Enacted statutes are public. Some texts were obtained from mirrors that publish under their own terms; the database carries the statutory words and the official source URL, not the mirror pages, and whether any mirror's terms reach downstream use of the transcriptions has not been assessed.

Do any export controls or other regulatory restrictions apply?

Not applicable.

7. Maintenance

Who will be supporting, hosting, and maintaining the dataset?

The AI Governance Collective, with Kyle David as editor.

How can the owner, curator, or manager of the dataset be contacted?

Through the AI Governance Collective at https://www.aigovcollective.com.

Is there an erratum?

Corrections are recorded per release in the dataset's PROVENANCE.md, which lists each audit's findings and the fixes applied, and every correction ships in the next numbered release. A public-facing erratum page separate from that file is not yet determined.

Will the dataset be updated? How often, by whom, and how will updates be communicated?

Yes. Legislative scans and maintenance passes run roughly monthly during session season; the next scheduled pass is October 1, 2026, after California's September 30 signing deadline. Each pass produces a new numbered release, and every page of the application shows the current release tag. Updates add new instruments, amend records when laws are amended, litigated, or take effect, and correct errors found by audit; excluded records are retained rather than deleted. A separate notification channel for members is not yet determined.

If the dataset relates to people, are there applicable limits on the retention of the data?

Not applicable. The only personal data is the public record of officials acting in office.

Will older versions of the dataset continue to be supported, hosted, and maintained?

Every release is a frozen, versioned snapshot kept in an append-only releases folder, so older versions are preserved. The application serves the latest release; whether members will be able to reach earlier releases through the application is not yet determined.

If others want to extend, augment, build on, or contribute to the dataset, is there a mechanism for them to do so?

Not at this release. Corrections and questions are welcome through the Collective, and anything accepted goes through the same adversarial fact-check as the project's own work before it enters a release. A formal process for external contributions is not yet determined.

How to cite this database

AI Governance Collective, US AI Law Database, release v0.6.0-2026-09-13, https://laws.aigovcollective.com, accessed <date>.

Questions and corrections

Send them to the AI Governance Collective at https://www.aigovcollective.com. Every correction that survives fact-check is logged in the release notes and shipped in the next numbered release.

Members can open the database at /laws. Not a member? Back to the landing page.