The Coder That Never Complains: A Full Accounting of JSON-LD Fidelity in CoreModels
Here is the fact that should make you suspicious: in either direction, the CoreModels JSON-LD coder emits no lossiness records at all. Not "few" — none. Most coders in the engine keep a list of things they apologize for; this one's ledger is empty. That is either an honest structural property or a place where losses hide, so this article is the accounting: what maps to what, what rides in the extras channel, what the `@id` contract guarantees, where fidelity breaks, and what the tests pin down.
The Coder That Never Complains: A Full Accounting of JSON-LD Fidelity in CoreModels
Here is the fact that should make you suspicious: in either direction, the CoreModels JSON-LD
coder emits no lossiness records at all. Not "few" — none. Most coders in the engine keep a list of
things they apologize for; this one's ledger is empty. That is either an honest structural property
or a place where losses hide, so this article is the accounting: what maps to what, what rides in
the extras channel, what the @id contract guarantees, where fidelity breaks, and what the tests
pin down.
We are the CoreModels team at ARAMAI.
Scope first. As a schema format, jsonld means a vocabulary: rdfs:Class and
rdf:Property nodes inside an @context/@graph document, the schema.org publication style.
Instance records go through a separate coder — the jsonld data format — where each @graph
node is a record keyed by @id, property keys are the elements' IRIs, references are
{"@id": "<record-id>"}, and a property the schema does not know is reported as a
StructuralDrop rather than thrown. That coder keeps its own ledger. The rest of this article is
the schema plane.
The structural mapping
One document that exercises the awkward corners — a shared property with two domains, a reference-typed property written with bare-string references, a date, and a node with no label:
{
"@context": {
"schema": "https://schema.org/",
"xsd": "http://www.w3.org/2001/XMLSchema#",
"ex": "https://example.org/shop#"
},
"@graph": [
{ "@id": "ex:Order", "@type": "rdfs:Class", "rdfs:label": "Order" },
{ "@id": "ex:Invoice", "@type": "rdfs:Class", "rdfs:label": "Invoice" },
{ "@id": "ex:placedOn", "@type": "rdf:Property", "rdfs:label": "placedOn",
"schema:domainIncludes": [ { "@id": "ex:Order" }, { "@id": "ex:Invoice" } ],
"schema:rangeIncludes": { "@id": "xsd:date" } },
{ "@id": "ex:reference", "@type": "rdf:Property",
"schema:domainIncludes": "ex:Order", "schema:rangeIncludes": "ex:Invoice" }
]
}
Decoding runs in two passes — classes and enumerations first, then properties and enumeration
members — so declaration order in the @graph is irrelevant. The mapping into our intermediate
representation:
| JSON-LD construct | IR construct | Detail |
|---|---|---|
node typed rdfs:Class | Type | label from rdfs:label, else the local name of the @id |
rdfs:subClassOf another class | inheritance | the first non-enumeration superclass wins |
rdfs:subClassOf schema:Enumeration | Taxonomy | lowercase schema:enumeration is accepted too |
node whose @type is an enumeration class | Taxonomy term | term id is the local name; the full @id is preserved |
node typed rdf:Property | Element | always decoded optional — RDF is open-world |
schema:domainIncludes | element-to-type wiring | every listed type gets the element |
schema:rangeIncludes | the element's value type | taxonomy ref, type ref, or primitive |
@id | mapsTo annotation | expanded against the context; types and elements also get provenance stamped JSON-LD with the @id as source ref |
@context | schema-level extras | every prefix binding preserved for re-emit |
For the document above that yields ex:Order holding ex:placedOn and ex:reference,
ex:Invoice holding ex:placedOn, a DateTime-typed placedOn, a reference typed as a pointer
to ex:Invoice, and an unlabeled property taking the label reference from its local name.
Bare strings, {"@id": …} objects, and arrays of either are all accepted anywhere a reference is
read.
Range resolution has a precedence worth knowing: a range naming a decoded taxonomy becomes a
taxonomy reference; one naming a decoded class becomes a type reference; an xsd: or schema:
datatype becomes a primitive; and a range pointing at an IRI the document never defines is still
a type reference — the open-world benefit of the doubt — rather than a string.
Datatypes, decoding: xsd:integer, xsd:int, xsd:long, schema:Integer → Integer;
xsd:decimal, xsd:double, xsd:float, schema:Number, schema:Float → Double; xsd:boolean,
schema:Boolean → Boolean; xsd:dateTime, xsd:date, schema:Date, schema:DateTime →
DateTime; anything else → String. Encoding: Integer → xsd:integer, Double → xsd:decimal,
Boolean → xsd:boolean, DateTime → xsd:dateTime, everything else → xsd:string. Four prefixes
are pre-loaded in both directions — rdf, rdfs, xsd, schema — so a document may use them
without declaring them.
The extras channel
The IR carries an open key-value bag on every node, and this coder uses three dotted jsonld.*
keys to move wire detail that has no structural home:
jsonld.context.<prefix>(schema level) — every prefix-to-namespace binding, the document's and the defaults. This is what lets the encoder re-emit your namespaces instead of inventing its own.jsonld.range(element level) — the originalrangeIncludesid. Datatype decoding is many-to-one, and this is the device that undoes it:xsd:datedecodes to DateTime, and on re-encode the preserved range beats the genericxsd:dateTime. Run the document above through decode and encode andex:placedOnreturns asxsd:date, exactly as authored, with both domains rebuilt.jsonld.id(taxonomy-term level) — the member node's full@id, so terms keep their IRIs even though the term id inside the IR is only the local name. A term that never saw JSON-LD has no such extra; its member id is minted by concatenation instead.
The @id contract, precisely
A vocabulary's identifiers are global, and identifiers are how the engine carries meaning across
standards. On decode, an @id written as a qname whose prefix is declared (or one of the four
defaults) is expanded and lifted into a mapsTo annotation: schema:name becomes
https://schema.org/name. That annotation reappears as x-maps-to in JSON Schema, as slot_uri
in LinkML, as a column comment in DDL — and a shipped test pins exactly that: a vocabulary pushed
through the engine into JSON Schema emits name with
"x-maps-to": { "schema": "https://schema.org/name" }.
On encode, each node's @id is chosen by a strict preference: an IR id that already looks like a
qname (a colon, no pipe, no space, no dot) is kept; otherwise the node's mapsTo IRIs are
compacted against the prefix map, longest namespace first; otherwise an id is minted as ex: plus
the PascalCased label.
Two consequences that matter more than they look:
Write qnames, not full IRIs. An @id of https://example.org/Device has a prefix of https,
which no context declares, so nothing is lifted into mapsTo — and on the way out it fails the
qname test (it contains dots), finds no mapsTo to compact, and is re-minted from its label as
ex:Device. Structure survives that trip; identity does not. If your vocabulary is authored with
full IRIs and you care about the identifiers, declare a prefix and use qnames.
A minted ex: id may have no binding. The emitted @context lists the prefixes used by the
document's identifiers that the coder can resolve. ex is not one of the four defaults, so a
model that never came from a JSON-LD document declaring it yields ex: qnames with no ex
binding. The vocabulary keys are the same story: rdfs:Class, rdf:Property, and
schema:domainIncludes rely on prefixes the coder always understands on the way back in, but a
general-purpose JSON-LD processor will want them declared. Merge the standard bindings in before
handing the file to one.
Where the honesty actually lives
The empty ledger is a structural outcome, not a cover-up: everything a vocabulary can express —
classes, subclassing, enumerations, domain- and range-typed properties, IRIs — has a first-class
home in the IR, and the three extras keys catch the wire details that do not. On a jsonld
transform, the entries in the response come from the other stages: the target encoder, the plan,
the gate.
That leaves three real limits, and one of them is quiet enough to deserve a warning:
- Required-ness disappears without a record. RDFS cannot say "this property must be present,"
so a required element encodes as an ordinary
rdf:Property— and no lossiness entry is produced. We measured it: a Customer model withcustomer_keyandfull_namerequired, encoded to JSON-LD and decoded back, returns four optional elements and an empty encode ledger. Inside pure JSON-LD round trips this is symmetric and harmless; across formats it is the first thing to check. - Constraints other than enumerations have nowhere to land. Max-lengths, patterns, numeric bounds, and cardinality bounds are not vocabulary constructs; enumerations are the exception, and travel both ways.
- Non-vocabulary nodes are ignored. A
@graphnode that is neither anrdfs:Class, anrdf:Property, nor a member of a decoded enumeration is skipped. Keep vocabulary documents to vocabulary content, and put instance data through the data plane.
A cosmetic asymmetry is also worth planning around when the model was not born in RDF. Encoding a
Customer type decoded from DDL, the type becomes ex:Customer and its properties
ex:CustomerKey, ex:FullName, ex:SignupDate, ex:Status — minted from the labels, because ids
like CustomerFullName are not qnames — while the taxonomy keeps its raw id and its members are
named by concatenation:
[
{ "@id": "CustomerStatusEnum", "@type": "rdfs:Class",
"rdfs:subClassOf": { "@id": "schema:Enumeration" },
"rdfs:label": "CustomerStatusEnum" },
{ "@id": "CustomerStatusEnumdraft", "@type": { "@id": "CustomerStatusEnum" },
"rdfs:label": "draft" }
]
The document re-decodes cleanly, but those ids are not dereferenceable IRIs. If you intend to
publish a vocabulary for other people to reference, give the governed model its mapsTo
identities first — then the encoder compacts real IRIs instead of minting placeholders.
Round trips and edge cases
Round-trip fidelity, as measured: decode → encode → decode preserves the structural signature —
types, parents, elements, value types — which is what the shipped test asserts; taxonomies with
their terms make the same trip. One cycle further (decode, encode, decode, encode) and the second and third documents are
byte-identical. The output converges rather than drifting, because the @context is sorted, the
@graph order is fixed (classes, then enumerations with members, then properties), and references
are normalized to {"@id": …} objects with domainIncludes always an array. The first hop is
structurally faithful rather than byte-faithful; from there it is stable.
The edges, from the code and from running it:
- A root that is not a JSON object fails immediately: path
$, message "The root of a JSON-LD document must be an object." - A root object with no
@graphis valid — it is treated as a single-node graph, so one free-standing class definition decodes. - Multiple
rangeIncludesvalues: the first drives typing and is the one preserved; the rest are not modeled. - Multiple
domainIncludesvalues: the element attaches to every named type, and on encode the domain array is rebuilt from every type that holds it — shared properties survive. - An element that belongs to no type encodes without a
domainIncludeskey;rangeIncludesis always written. - Missing
rdfs:labeldegrades to the local name after the last:or/.
That is the whole accounting. The coder itself has nothing to apologize for, because identifiers — not field names — are what the engine carries between standards. What a vocabulary cannot carry is required-ness and value constraints; check those two, and the rest travels. For the surfaces that drive this coder, see the transform section of the CoreModels docs.