Projections: one record, every framework
A Governed Data Record carries its own definitions. Every catalog entry that describes it is generated from it, in that community's format, and checked with that community's own validator. The record never changes.
What this is, and what it is not
Dataset-description frameworks ask a publisher for the same things: what the dataset is, what each variable means, which code lists it uses, how absence is marked, where the schema is, who to contact. CDIF asks for them in JSON-LD with SHACL shapes; the United States federal catalog asks in a JSON Schema called data.json; Europe asks in RDF under DCAT-AP, and its health data space adds a data dictionary under HealthDCAT-AP. Each is a top-down profile: it tells a publisher how to describe data it already holds.
A model published on the Semantic Data Charter already holds every one of those answers, per variable, inside the record: the datatype, the unit, the value domain, the sixteen exceptional values for missing and not-applicable, the schema by identifier and hash, the terminology bindings. So a description is not written; it is projected. One reader opens the published package, one writer emits the target format, and the target community's own validator says whether the result conforms. We hold the record still and let four judges look at it.
What it is not: a claim that these frameworks are wrong, or that a projection says everything a framework can say. Each repository's README has a section on what the projection could not say, left out rather than filled in, and a section on what we learned about the profile. No issue was filed on any of these projects; the findings are recorded on our side.
The record in every row below is the same one: the NHANES Participant model from the FAIR Data Demo (153 variables, 67 code lists), read from its published package by sdcreader. Every validator is pinned to the commit it was run at, so a passing result records which version of the framework it passed.
The four, each with its own judge
CDIF
CODATA's Cross-Domain Interoperability Framework, handbook 1.1
The projection: a JSON-LD document with the dataset, 153 instance variables each carrying its datatype, unit and role, 67 concept schemes for the code lists, and one sentinel value domain for the sixteen exceptional values.
The judge: CDIF's own conformance validator, run on the framework's SHACL rules at a pinned commit, and the codelist building block. Zero violations; the declared profiles agree with the detected ones.
Could not say: a distribution of the records (the model is public; the records are not), temporal coverage, statistics; all recommended, none required.
sdc_cdif on GitHubDCAT-US 3.0
The United States federal data.json standard (GSA)
The projection: a data.json catalog whose data-dictionary slot, the one OMB M-25-05 requires for "the names and definitions of all variables", holds the published schema itself, pinned by hash. Seven FAIR Data Demo models as one catalog.
The judge: GSA's JSON Schema at a pinned commit, and Data.gov's online validator. No errors, for the one model and for the seven.
Could not say: anything an agency would have to determine (open-data, FOIA and licensing determinations, bureau and program codes); taken only as declared input, never defaulted.
sdc_dcat3_us on GitHubDCAT-AP 3.0.1
The European application profile for data portals (SEMIC)
The projection: a catalog in Turtle and JSON-LD, every referenced node typed and named, the schema as the dataset's standard, the European vocabularies for theme and access rights.
The judge: SEMIC's SHACL shapes offline and the Interoperability Test Bed online, its four validation groups each conforming. One finding on the way: the Test Bed's combined type reports eleven violations an empty catalog receives too, from its own background knowledge; the groups are the honest judge.
Could not say: distributions, temporal coverage, geometry, frequency, series, data services; and the things a publisher decides (theme, access rights, language), declared and said so.
sdc_dcat3_ap on GitHubHealthDCAT-AP
The European Health Data Space profile for secondary use, release 8
The projection: the first target with a real data dictionary: every leaf of the record is a column of the dataset's CSVW table group, and the model's SNOMED CT, LOINC and NCIT bindings become its coding systems.
The judge: the profile's own shape set for non-public data, zero violations, and the Test Bed engine on the same shapes agreeing. One finding: with DCAT loaded the validator reads a catalog as a dataset and holds it to the health mandatory list, which is why the profile's own examples are dataset-only.
Declared, not read: the health data access body, the applicable legislation, the population and personal-data statements; the holder's facts, said so in the README.
sdc_healthdcat_ap on GitHubVerify it yourself
Each repository pins the framework's validation artifacts at a commit under data/, so the check runs offline and says which version it passed. The README's second section is the command. The pattern is the same four times: install the package, point it at the published model package (or let it fetch the package from the catalog by identifier), write the description, run the framework's own validator over it.
Two kinds of input are kept apart on purpose. What comes from the package is never invented: the variables, the code lists, the schema and its hash, the model's own Dublin Core when the modeler wrote it. What a publisher has to decide (a contact point, a theme, access rights, a language, who grants access) is declared input, named as such in each README's third section, and defaulted only for the sample.
Next in the sequence: schema.org Dataset and Croissant. The same record, two more judges.
One foundation, every framework a projection. The same move as a partner's mandated document format or a FHIR resource: the Governed Data Record is the primitive, and the descriptions, documents and graphs are projections of it.