FBI TSC AI Sources Sought: What “Traceable Lineage” Actually Requires
Date: August 16, 2026 · Author: Dmitrii Zatona · Updated: September 29, 2026
TL;DR
- The FBI Threat Screening Center posted an AI sources sought notice, FBI-TSC-AIE, on March 27, 2026; it is archived, and no public follow-on solicitation or award appears as of September 29, 2026 (Section 1).
- Read as engineering, its predictive-modeling use case describes record linkage across federated systems: a prediction of where data about an entity may reside, which is a different task from predicting behavior (Section 2).
- Shipping lineage catalogs record dataset and column edges in mutable stores, research systems record provenance per row without integrity protection, and verifiable logs exist for certificates, software and media (Section 3).
- Derived attributes can return through training data as apparent corroboration, and corrections cannot reach outputs through a record-level graph that was never stored (Section 4).
- Under deliberately favorable assumptions, most flags from a per-person score against a rare outcome are false; an outcome evaluation of one policing score’s first version found no measured benefit, and a second review found records too poor to judge (Sections 5 and 6).
- The analysis does not claim the TSC will build a risk score or that any vendor offered one; its negatives are bounded to the documents named, as of September 29, 2026 (Section 8).
Traceable lineage is the phrase that turns an FBI market-research notice into an engineering problem. The Threat Screening Center asked vendors how they would build “Predictive Modeling Using Enhanced Data with Traceable Lineage,” and the text behind that title asks for traceability to the specific source data, enhancement logic and reference materials behind every prediction. Taken at face value, that is reconstructible provenance for each individual output, kept in a form someone who doubts it can check.
The phrase reads like standard procurement language, and the distance between it and the tools the data industry ships is the subject here. This article takes the requirement apart into the record-linkage task it describes and the five provenance problems that task inherits.
1. What was published
The document is a sources sought notice, identifier FBI-TSC-AIE, titled “Threat Screening Center AI Enhancement.” The FBI posted it on SAM.gov on March 27, 2026, with responses due April 10, 2026 and a place of performance in Clarksburg, West Virginia. The attached requirement document, labeled Attachment AS, version 1.2, and issued by the TSC’s Information Technology Unit, is publicly downloadable from the notice’s attachments.
A sources sought notice is market research. It is not a solicitation, a request for proposal or a contract, and the document says so in its disclaimer. Its introduction adds that responses may be used to formulate requirements and acquisition strategies for competitive solicitations. Vendors were asked for capability statements of up to 25 pages and whether they hold a GSA Schedule contract.
The notice was archived on April 25, 2026 and republished on July 22, 2026 with one substantive change in its background section, discussed in Section 2; the republished version is archived as well. As of September 29, 2026, an exact-phrase search of SAM.gov for “threat screening center” returns no follow-on solicitation, and USAspending shows no award traceable to the notice. That is a statement about public records. Above the simplified acquisition threshold, orders competed among GSA Schedule holders are either posted on GSA’s eBuy or sent directly to Schedule contractors (FAR 8.405-2 , paragraph (c)(3)(iii)), and none of the TSC’s analysis-services orders described in Section 7 appears on SAM.gov as a public solicitation.
The notice drew press attention after Reason reported on it on July 28, 2026 (Daniel Boguslaw).
2. What the requirement says
The RFI lists six use cases, and every one of them carries an explicit traceability or citation clause. Five are retrieval and reporting: an AI knowledge base over internal policies and SOPs with mandatory source citations; federated search across repositories with source attribution; identifier-driven recurring reports with per-element citations; natural-language query with synthesized answers that distinguish extracted facts from AI-generated synthesis; and natural-language-driven data visualization with citation metadata. The remaining one, use case 5, is the subject here. In full:
“The solution must leverage existing enterprise datasets that include enriched or enhanced data elements with documented source attribution to develop predictive models. When new data is ingested, the system must analyze similarity, pattern alignment, and attribute correlation against existing records to predict where additional relevant information may be derived across federated systems. The solution must provide transparent reasoning for its predictions, including traceability to the specific source data, enhancement logic, and reference materials used in generating the inference. All predictive outputs must retain auditable lineage to ensure analytic defensibility and support oversight requirements.”
Read as an engineering specification, this decomposes into four known problems.
Analyze similarity, pattern alignment, and attribute correlation against existing records is probabilistic record linkage: deciding whether two records that are not textually identical refer to the same entity. The canonical framework is Fellegi and Sunter’s “A Theory for Record Linkage” (JASA 64(328), 1969), which scores agreement patterns between record pairs and sorts them into link, possible link and non-link. Open implementations include Splink , from the UK Ministry of Justice, Zingg and dedupe ; Senzing is the commercial reference point.
Predict where additional relevant information may be derived across federated systems is candidate generation and link prediction: given a new record, rank which external systems and records are likely to hold related information. In record-linkage terms this is blocking, which prunes the pair space to the candidates a comparison should consider (Christen, IEEE TKDE 24(9), 2012).
This matters because it is a different task from predicting a person’s future behavior. Reason characterized the goal as AI to “help predict who might be a terrorist.” The requirement’s text describes a system that predicts where data about an entity may reside, and it says nothing about what the entity will do. Both tasks are hard, and they fail differently. A record-linkage system’s characteristic failure is misidentification, merging two people into one identity or scattering one person across several; a behavioral-prediction system’s characteristic failure is the base-rate problem quantified in Section 5.
A critique of the requirement as behavior prediction therefore critiques a system the text does not describe. The distinction collapses only if linkage scores are later repurposed as risk scores, and the document neither states nor precludes that.
Enriched or enhanced data elements are derived attributes: values computed by earlier processing, such as a resolved identity, a normalized name or a model-assigned label, that exist in no source system. Section 4 shows why they dominate the difficulty.
Documented source attribution and auditable lineage name provenance. The distance between that phrase and what shipping tools do is the subject of Sections 3 and 4.
One more clause matters for how the notice has been read. The RFI’s background section grounds the TSC’s mission in Homeland Security Presidential Directive 6 and National Security Presidential Memorandum 7. Two unrelated documents carry that designation, each self-designated NSPM-7 in its official text. The first is the memorandum of October 4, 2017, “Integration, Sharing, and Use of National Security Threat Actor Information to Protect Americans” , the directive behind watchlisting-data interoperability. The second is the memorandum of September 25, 2025, “Countering Domestic Terrorism and Organized Political Violence” (90 FR 47225).
The original March 27 PDF cited NSPM-7 without a date. The July 22, 2026 republication resolved the ambiguity. Its amended background dates the reference to October 5, 2017, the day the White House released the 2017 memorandum, and names the threat-actor-information memorandum. The Reason article, published July 28, described the reference as likely connected to the 2017 directive; the amended notice had been posted on SAM.gov six days earlier.
3. Lineage versus provenance
The data industry uses “lineage” and “provenance” close to interchangeably. For this requirement the difference is the entire problem, so the terms need separating.
Lineage, as shipped, is a graph over data assets: this table feeds that job, which writes that table. The emerging standard, OpenLineage , models jobs, runs and datasets. In version 1.53.0 of the specification, its finest standardized mapping runs field to field: the column-level lineage facet and the lineage facet that supersedes it map input columns to output columns. Its subset facets can describe which part of a dataset a job read or wrote, by location, partition or a condition on field values, but not which input rows produced which output row. OpenLineage’s reference implementation, Marquez , stores these events in PostgreSQL. The catalog DataHub records lineage down to the column and accepts it by automatic extraction, by API or by manual editing in its interface; its documentation warns that “lineage added by hand and programmatically may conflict with one another.”
Apache Atlas models process entities whose inputs and outputs are arrays of datasets. Databricks Unity Catalog captures lineage down to the column level. In each of these systems as documented, the built-in lineage model bottoms out at column granularity, usually at the dataset, and the metadata store is an ordinary mutable database recording what pipelines assert about themselves.
Provenance, in the sense this RFI needs, is a statement about an individual output record: these input records, transformed by this code at this version, produced this value, retained in a form that supports later scrutiny. The W3C PROV data model (Recommendation, 2013) can express that statement, because its entities can be defined at any granularity and wasDerivedFrom can link individual record-entities. But PROV is a vocabulary, and a vocabulary enforces nothing. It does not require the graph to be complete or protect it from modification; the companion note on access, PROV-AQ , addresses tampering only as informal guidance, with HTTPS as its concrete recommendation.
Record-level provenance is not unknown to computer science. Provenance semirings (Green, Karvounarakis and Tannen, PODS 2007 ) annotate every tuple with a polynomial recording how it was derived. The ProvSQL extension implements them in PostgreSQL, Trio and GProM are earlier systems in the same line, and the survey by Cheney, Chiticariu and Tan (Foundations and Trends in Databases 1(4)) codifies the theory.
The closest industry system is Pachyderm, whose documentation describes commit provenance: an output commit records the input commits it was derived from. Its versioned unit is the commit of files in a repository, and its immutability is operational rather than tamper-evident. Cryptographically protected provenance chains exist in research (Hasan, Sion and Winslett, FAST 2009 ) and existed briefly as a managed cloud product. Amazon QLDB, a ledger database with a cryptographically verifiable transaction log, reached its end of support on July 31, 2025.
The state of the art therefore splits three ways. Record-level granularity exists in research systems without integrity protection; integrity protection exists in systems without pipeline lineage; industry lineage catalogs have neither. Among the tools and specifications in this section, as documented on September 29, 2026, none combines record-level granularity with tamper evidence, and that combination is what auditable lineage in service of analytic defensibility of individual predictions requires.
A lineage catalog’s answer to the question why is this attribute on this record is an assertion by the system being questioned. The difference between an assertion and evidence surfaces at one moment: when someone disputes the assertion.
4. Five provenance problems the requirement inherits
Each problem below follows from the gap in Section 3 and from one clause of use case 5. They run from the single pipeline outward: capture at the record level, feedback through derived attributes, correction after the fact, reconstruction of a past decision, and verification between operators who do not trust each other.
4.1 Provenance survives transformation only if captured at the record level
A pipeline joins records from systems A and B, applies a model and writes an output record. A lineage catalog records that job J read datasets A and B and wrote dataset C, possibly with column mappings. It does not record which rows of A and B produced which row of C, under which code version, with which parameters. Documented source attribution implemented on catalog metadata attributes the output to systems rather than to statements: it can say a value came from somewhere in dataset A, but not which record of A, as of when, under what transformation.
Federal control language already names the missing properties. In NIST SP 800-53 Rev. 5 , control AU-9 requires protecting audit information from modification and deletion, and enhancement AU-9(3) adds cryptographic mechanisms to protect its integrity. AU-10 requires non-repudiation: irrefutable evidence that a given actor performed a given action. A mutable metadata store populated by self-reported pipeline events does not, by itself, provide either property.
The catalogs are not defective. They were built for impact analysis and debugging, where a trusted operator is a reasonable assumption. The RFI’s phrase analytic defensibility describes an adversarial setting, and the assumption does not transfer.
4.2 Derived attributes contaminate their own future evidence
The requirement says the enterprise datasets already contain enriched and enhanced elements and that new predictive models are to be developed on top of them. That creates a loop.
The diagram is an inference from use case 5’s description of models developed on enriched datasets, not something the RFI defines. An enrichment model computes a derived attribute, say a resolved-identity link, and the attribute is written into the consolidated record. The consolidated store is later sampled to train the next model, whose outputs now partially encode the first model’s guesses.
When the second model emits an attribute that agrees with the first, the agreement looks like independent corroboration. It is the same inference, recycled through a training set. The provenance graph of the record store has stopped being acyclic, and any process that counts how many indicators support a link, human or automated, can count one indicator several times.
Detecting the cycle requires the record-level edges that Section 4.1 showed are not captured. At dataset granularity the loop is invisible: “enrichment dataset feeds training dataset feeds enrichment dataset” describes both a healthy retraining loop and a self-confirming one. Only the record-level graph distinguishes them, and none of the catalogs in Section 3 stores it.
The screening context gives a concrete reason to care about compounding derived signals. The Fourth Circuit’s opinion in Elhady v. Kable, 993 F.3d 208 (2021), recites from the factual record that the TSC received about 113,000 nominations a year and accepted about 99% of them. At that acceptance rate, the provenance of attributes supplied with nominations, computed ones included, is material to what the consolidated store contains.
4.3 Corrections do not propagate through a graph nobody stored
Suppose a source record is corrected, or a nomination is withdrawn. Every derived attribute computed from that record, every model trained on those attributes and every prediction served from those models is now partly built on retracted input. The technical name for handling this is revocation propagation: walk the transitive closure of everything downstream of the retracted fact and invalidate or recompute it. The cost of a correction is the cost of that walk. If the record-level graph was never stored, the walk cannot be performed, only approximated, typically by full recomputation or by not propagating at all.
None of the lineage catalogs in Section 3 defines revocation semantics. Their edges say where data came from; they cannot say that an input is no longer valid and what must be recomputed. The Privacy Act of 1974 anticipates the need at the record level. Under 5 U.S.C. § 552a(c) , each agency must keep an accounting of the date, nature and purpose of each disclosure of a record and of its recipient, except disclosures to its own employees who need the record for their duties and disclosures required under the Freedom of Information Act. It must retain that accounting for at least five years or the life of the record, and under (c)(4) inform prior recipients of later corrections and disputes.
The FBI’s system-of-records notice for TSC records (JUSTICE/FBI-019, 76 FR 77846 ) claims exemptions under § 552a(j) and (k), implemented at 28 CFR 16.96(r) : the system is exempt from (c)(3), the subject’s access to the accounting, and from (c)(4), correction notices to prior recipients, among others. Subsection (j) cannot exempt a system from (c)(1) and (c)(2), the duty to create and retain the accounting. The engineering consequence mirrors the legal one: the disclosure-and-derivation graph must exist and be maintained regardless of who is entitled to read it.
The operational numbers give the propagation problem scale. A GAO report, GAO-25-108349 (August 2025), counts roughly 20,000 redress inquiries that U.S. persons submitted to DHS TRIP from December 2021 through September 2023. Of those, 1.5% (289) related to the terrorist watchlist, and about a third of the watchlist-related inquiries (88) ended with the person removed from it. A DOJ OIG audit, 14-16 (2014), measured a median of 78 days for the FBI to remove certain non-investigative subjects from the watchlist. A removal is a revocation event. Whether it reaches every derived attribute and model that consumed the record depends on whether the consumer graph exists.
The retention schedule sets the time horizon. The FBI-019 notice keeps active screening records for 99 years, archived ones for 50 and audit logs for 25. Any lineage scheme adopted for this system therefore has to stay interpretable across several generations of software, formats and vendors, which favors simple, self-describing, verifiable structures over any one product’s internal metadata schema.
4.4 Reproducing a decision at time T requires bitemporal, immutable state
Transparent reasoning for its predictions is a statement about the past by the time anyone asks for it. To reconstruct why the system produced a given output years earlier, four things must be recoverable as they were at that moment: the feature values, the model version, the thresholds and rules, and the mapping between them. This is the bitemporal problem, distinguishing valid time (when a fact was true in the world) from transaction time (when the database recorded it). Snodgrass formalized it in Developing Time-Oriented Database Applications in SQL, and modern bitemporal SQL systems such as XTDB implement it.
The nearest industrial practice is the feature store. Both Feast and Tecton perform point-in-time-correct joins, so training data contains only feature values that were available at each historical event’s timestamp. This is the right machinery, but its guarantee is prospective: it prevents leakage when training sets are built. It does not make the historical state provable afterward, because the join is only as good as the mutable offline store it reads at query time.
Adjacent versioning tools stop at the same line. Delta Lake time travel reaches old table versions only until VACUUM runs, and by default that command removes unreferenced files older than seven days. The MLflow registry versions models; DVC and lakeFS version files by content hash. All of them detect change. None of them produces evidence that history was not rewritten, because their storage remains administratively mutable and their references movable.
Litigation over the No Fly List shows what reconstruction demands in practice. In Latif v. Holder, 28 F. Supp. 3d 1134 (D. Or. 2014), judicial review ran on an administrative record containing the information the government relied on for a listing decision, furnished to the court and not to the petitioner. The revised process in the 2016 follow-up opinion turned on whether a statement of reasons and the underlying material could be produced and reviewed.
In FBI v. Fikre, 601 U.S. 234 (2024), a government declaration that a removed individual would not be relisted on the basis of currently available information failed to moot the case. As the Supreme Court summarized the Ninth Circuit’s reasoning, the declaration “does not disclose what conduct landed Mr. Fikre on the No Fly List,” and the Supreme Court found it too sparse to show that the listing could not reasonably be expected to recur. Both cases reduce, on the technical side, to the same artifact: a reconstructible account of which inputs produced a determination, at a specific time. Section 4.1’s record-level provenance and this section’s bitemporal snapshots exist to produce that artifact.
4.5 Across a trust boundary, an assertion is not evidence
The notice does not define federated systems; what the record shows is distribution across operators. Per the factual record recited in Elhady, the TSC makes the screening database available to federal, state and tribal law enforcement agencies, some foreign governments and some private companies working in sensitive security areas. The notice’s own summary describes querying external systems for data enrichment.
Where record data or its provenance metadata crosses between operators, the receiving side can check the claimed history only if a verification mechanism exists. None of the catalogs in Section 3 supplies one, so the receiving side can only store the claim. Every problem earlier in this section then compounds: the record-level graph is split across operators, each part mutable and none independently checkable.
This shape, many mutually distrusting parties relying on each other’s claims about history, has been addressed in three domains with a recurring pair of primitives.
Certificate Transparency (RFC 6962 , RFC 9162 ) made certificate issuance auditable. Every certificate goes into an append-only Merkle-tree log. An inclusion proof shows that a given entry is in the log, and a consistency proof shows that the log at time T₂ extends the log at T₁ without rewriting it, so independent monitors can check history continuously. Sigstore’s Rekor applies the same structure to signed supply-chain metadata.
The in-toto attestation format binds signed metadata to artifacts by cryptographic digest. On top of it, SLSA v1.0 layers build-provenance levels, up to builds on a hardened platform where forging provenance requires exploiting a vulnerability beyond the capabilities of most adversaries. The C2PA specification does the same for media: signed manifests whose hash bindings break if the asset’s bytes change.
None of these domains settled on a richer metadata catalog as the working answer. Two primitives recur. The first is metadata bound to the individual artifact by digest and signature: the in-toto, SLSA and C2PA layer. The second, wherever parties must audit one another’s history over time, is a verifiable append-only log of those bound statements, checkable by parties who do not trust its operator: the CT and Rekor layer. Translated to this RFI’s setting, entries would bind to individual records and their derivations, and proofs would show that the derivation history at decision time is a prefix of the history at review time.
The diagram’s right side is an inference from how Certificate Transparency and Rekor work, not something the RFI or any lineage standard defines. The regulatory direction of travel is compatible with that reading and stops short of mandating it. OMB M-25-21 (April 2025), which rescinded and replaced M-24-10, presumes law-enforcement risk assessments about individuals and identification of criminal suspects to be high-impact AI. For such uses it requires pre-deployment testing, an impact assessment, ongoing monitoring with processes enabling traceability where possible, and human oversight, and it excludes AI used as a component of a national security system.
The EU AI Act (Regulation 2024/1689 ) requires high-risk systems to record events automatically over their lifetime (Article 12), with providers and deployers keeping the logs for at least six months (Articles 19 and 26(6)). It classifies specified law-enforcement uses, including individual risk assessment and profiling, as high-risk (Annex III, point 6). Its 2026 “Digital Omnibus on AI” amendment, Regulation 2026/1744 , left those texts intact and postponed their application for Annex III systems to December 2, 2027.
Both regimes mandate that logs and documentation exist. Neither specifies the property this section is about, that the logs be verifiable by a party who does not trust their keeper. NIST’s AI RMF (AI 100-1) names maintaining the provenance of training data as an aid to transparency and accountability and leaves the mechanism open.
5. Scale
One quantitative constraint applies to any system that scores rare events. It needs explicit assumptions, because it is routinely misattributed to model quality.
Axelsson’s analysis of intrusion detection (CCS 1999 ) formalized it. When the condition being detected is rare, the probability that an alarm is genuine is dominated by the false-positive rate rather than by the detection rate; in his words, “the false alarm rate is the limiting factor.” The arithmetic is three lines of Bayes. Let p be the fraction of screened items that are true positives, d the detection rate and f the false-positive rate. The probability that a flagged item is a true positive is pd / (pd + (1−p)f).
Take a deliberately favorable hypothetical: a screening population of 1,000,000, of whom 100 are true positives (p = 10⁻⁴), a detector with d = 0.99, and f = 0.01. A 1% false-positive rate is strong for models over noisy, heterogeneous records. Flagged items then number 99 true and 9,999 false, and the probability that a given flag is genuine is about 0.98%. Cutting f tenfold, to 10⁻³, raises it only to about 9%. For the posterior to reach even two-thirds, f must approach p: for this population, a false-positive rate near 5×10⁻⁵, two orders of magnitude below the favorable assumption. These are properties of the base rate rather than of any particular model, and they apply identically to medical screening for rare conditions.
Section 2’s distinction matters here. The requirement as written is record linkage, whose output is “look in system X,” and linkage precision is measurable and improvable against ground truth. Any layer that converts federated-match signals into a per-person score against a rare outcome inherits this arithmetic in full, multiplied by the record volumes involved: the Elhady record puts the screening database at 1,160,000 individuals. The numbers above say what the flag stream of any such layer would be made of. They say nothing about whether to build one, which is not an engineering question.
6. What evaluation has shown elsewhere
Two person-based predictive scoring programs in U.S. policing, Chicago’s and Los Angeles’s, have been reviewed with access to internal data.
RAND researchers evaluated Chicago’s Strategic Subject List (Saunders, Hunt and Hollywood, Journal of Experimental Criminology 12(3), 2016) with a quasi-experimental design on version 1 of the model, which listed 426 people in 2013. People on the list were no more or less likely to become homicide or shooting victims than a matched comparison group, and the model identified fewer than 1% of homicide victims in its window (3 of 405). They were 2.88 times more likely to be arrested for a shooting. The authors’ mediation analysis found that this effect did not run through additional police contact, and they suggest officers may have used the list as leads for closing shooting cases.
The City of Chicago OIG’s 2020 advisory added the operational record. As of July 2018, 399,412 people had SSL risk scores; the scores went un-updated from August 2016 until January 2019; neither CPD nor RAND evaluated versions 2 through 5 of the models; and CPD decommissioned the program on November 1, 2019.
The LAPD Inspector General reviewed the chronic-offender part of Operation LASER in March 2019 . Of the people in the chronic-offender database with detailed point calculations, 44% had zero or one violent-crime arrest, and about half had no gun-related arrests. Site visits found the program applied inconsistently across areas, and the OIG identified significant barriers to evaluating it, the most significant a “lack of clear, reliable data” on both inputs and outcomes. The department had suspended the program and its tracking database in August 2018.
Two programs do not generalize to all predictive systems. Their findings are of two kinds. RAND’s outcome evaluation measured no effect on the harm the program targeted, for version 1 of the list. The operational reviews documented a different problem: Chicago’s later model versions were not evaluated by CPD or RAND and ran on unreliable, static scores, and the LAPD OIG found the chronic-offender records too unreliable to measure the program’s inputs and outcomes. The second kind is a lineage finding.
7. Who would build it
The RFI does not arrive into empty space. The TSC already buys analysis services at scale, and the notice describes an automation layer over data those services already work with. Two delivery orders for TSC Intelligence Analysis Services define the landscape, both competed under multiple-award GSA Schedules with seven offers received, per USAspending and FPDS records retrieved on September 29, 2026.
BAE Systems Technology Solutions & Services holds order 15F06725F0001209 , signed September 2, 2025, with an initial value including options of $128.7M. Three modifications reduced it to $122.0M, the last signed March 13, 2026; its base period ran to September 10, 2026, with options to September 2030. BAE also held a bridge order for the same services, 15F06725F0001838, signed September 30, 2025 with one offer received and a completion date no later than May 31, 2026.
IntelliWare Systems Inc. (parent: IntelliBridge LLC) holds order 15F06726F0000362 for the same services, signed March 13, 2026. Modifications later that month brought its value including options to $144.7M, and one signed September 21, 2026 reduced it to $142.9M; its potential period runs to October 2031.
The two main orders were competed through the Schedule fair-opportunity process, and none of the three appears on SAM.gov as a public solicitation, consistent with the absence of public follow-on records for the RFI noted in Section 1. Whatever becomes of this market research, the plausible path runs through the same contracting channel, into an ecosystem that is already occupied.
8. What does not exist as a standard
The table below sets each property the requirement implies against the closest existing tooling and what that tooling lacks.
| Property the requirement implies | Closest existing tooling | What it lacks |
|---|---|---|
| Record-level derivation graph | Provenance semirings, ProvSQL (research) | Standardization, integrity protection |
| Pipeline lineage capture | OpenLineage, DataHub, Atlas, Marquez | Record granularity; store is mutable and assertion-based |
| Tamper-evident history | CT (RFC 9162), Rekor, in-toto/SLSA, C2PA | Applies to certificates, software, media; not to data-pipeline records |
| Point-in-time state | Feature stores, bitemporal DBs, Delta time travel | Prospective correctness only; history administratively deletable |
| Revocation propagation | None among the tools surveyed | Not defined in any tool surveyed here |
| Cross-boundary verifiability | Transparency-log monitors and auditors | Not defined for federated record systems |
Five of the six properties have partial tooling somewhere; revocation propagation has none among the tools surveyed, and no single tool reviewed covers all six. One managed product that combined verifiable history with a database interface, Amazon QLDB, reached end of support in 2025.
In the documents reviewed here (the OpenLineage specification 1.53.0 with its column-lineage, lineage and subset facets, DataHub, Apache Atlas 2.0, Unity Catalog, W3C PROV-DM and PROV-AQ, Pachyderm, RFC 9162, Rekor, in-toto, SLSA v1.0 and C2PA 2.4, as of September 29, 2026), none defines per-record cryptographic bindings anchored in an append-only, independently verifiable provenance log for records in ML pipelines. That is the combination Predictive Modeling Using Enhanced Data with Traceable Lineage requires when implemented to the standard its own wording sets: traceability to specific source data and enhancement logic, auditable lineage on every predictive output, across federated systems, in a records system whose retention schedule runs to 99 years. The industry has built it for certificates and software artifacts, and in signed-manifest form for media files.
The same gap between an assertion and evidence appears when the output is generated text. The INL Reviewing Smarter RFI teardown reads a nuclear-licensing notice’s “every output traceable to source documents” as four properties of increasing strength, and among the tools surveyed there only the weakest of them ships.
The public record, as of September 29, 2026, does not show whether any response to this sources sought notice proposed such a primitive, or whether an eventual procurement will hold vendors to the sentence as written. This analysis does not claim that the TSC intends to build a per-person risk score, that any vendor offered one, or that the documents reviewed here exhaust the field. I work on verifiable provenance for records and the data derived from them, and that work is separate from this analysis. More write-ups are on the engineering blog.
If a record system has to carry its own derivation history, so that a correction reaches every attribute and model built on the old value, designing that layer is contract work I take on.