Apache Avro logoDeep dive

Every Avro Construct, and Where It Lands in the IR

Does an `int` come back as an `int`?

Every Avro Construct, and Where It Lands in the IR

Does an int come back as an int?

That question decides whether a schema converter is a tool you run or a tool you trust. CoreModels routes every format through one neutral intermediate representation (IR), and an IR that also hosts JSON Schema, SQL, LinkML, OWL, and Protobuf cannot have a distinct primitive kind for each of Avro's numeric encodings. So the honest answer has three parts: what maps structurally, what rides alongside as preserved fact, and what is declared as loss. This is the complete account for Avro, in that order.

The structural mapping

Import is decode (Avro → IR); export is encode (IR → Avro). The correspondences:

Avro constructIR constructNotes
recordType (id = record name)the document root must be a record
fieldElement (id Record.field)a plain field decodes as required — in Avro it always carries a value
enumTaxonomy (id = enum name)symbols become terms, in order, 1:1
["null", T] unionoptional Elementencode re-emits optionality as ["null", T]
arraycollection cardinality on the Elementthe item type becomes the Element's value type
nested recordType + type referencedefined inline at first use on encode
named-type referencetype reference to the same Typeshared references resolve to one Type
union of 2+ non-null branchesthe first branchdeclared as narrowing
map, fixedString approximationdeclared

Primitives decode coarsely: int and long → Integer; float and double → Double; boolean → Boolean; string, bytes, and null → String. Logical types refine that: date, timestamp-millis, timestamp-micros, and time-millis all decode as DateTime; decimal decodes as Double with a declared approximation, because the IR does not model precision and scale.

Encoding IR that never saw Avro, the defaults are Integer → long, Double → double, Boolean → boolean, everything else → string, and DateTime → {"type": "long", "logicalType": "timestamp-millis"}.

The extras channel

If coarse decoding were the whole story, every round trip would widen types. It is not the whole story. Every IR node carries an annotation bag, and the Avro coder writes the facts the IR cannot express into it under dotted keys:

KeyWhereWhat it holds
avro.typeElementthe exact original type token (int, float, bytes, or the logical type name)
avro.namespaceTypethe record's namespace
avro.docType, Elementdoc strings
avro.defaultElementthe field's default, as JSON text
_avroRootschemawhich record was the document root

The encoder reads them back, and the consequences are concrete. A field that arrived as int re-emits as int; float stays float; bytes stays bytes. A DateTime re-emits its own logical type on the correct physical base — date and time-millis on int, timestamp-millis and timestamp-micros on long. Namespace, docs, and non-null defaults reappear where they were. So this record:

{
  "type": "record",
  "name": "Shift",
  "fields": [
    { "name": "onDate", "type": { "type": "int", "logicalType": "date" } },
    { "name": "startsAt", "type": { "type": "int", "logicalType": "time-millis" } },
    { "name": "micros", "type": { "type": "long", "logicalType": "timestamp-micros" } },
    { "name": "bytesField", "type": "bytes" },
    { "name": "floatField", "type": "float" },
    { "name": "intField", "type": "int" }
  ]
}

decodes to three DateTimes, a String, a Double, and an Integer — and re-encodes to exactly the JSON above, token for token. One caveat lives here: a logical type the coder does not model (uuid, say) decodes as its base type and is not re-emitted, and nothing in the ledger flags it. If you rely on uuid semantics, check for it yourself.

Named types get the treatment the specification expects. Given a record whose billing field defines an inline Address and whose shipping field references "Address" by name, the decoder resolves both to one Type, and the encoder writes Address inline at first use and the bare name afterwards — no duplicated definitions, no dangling references. Optionality composes with all of it: a ["null", {"type": "array", "items": "string"}] field decodes as an optional collection and re-emits in exactly that nested shape.

Semantics without IRIs

Avro has no IRIs, so there is nothing like a predicate URI to hang meaning on. The escape hatch is that the specification tolerates unknown attributes in schema JSON, and the coder uses exactly that. Four attributes are lifted from records and fields into the IR's semantic carrier and re-emitted on encode: x-sia-role, x-sia-priority, x-sia-instruction, and x-maps-to. A record annotated like this:

{
  "type": "record",
  "name": "Patient",
  "namespace": "org.example.clinical",
  "x-sia-role": "master-data",
  "fields": [
    { "name": "mrn", "type": "string",
      "x-sia-role": "identifier",
      "x-sia-priority": 1,
      "x-sia-instruction": "Never leaves the trusted zone.",
      "x-maps-to": { "hl7": "http://hl7.org/fhir/Patient.identifier" } },
    { "name": "birthDate", "type": { "type": "int", "logicalType": "date" } }
  ]
}

decodes into a Type carrying the role master-data and an Element carrying role, priority, instruction, and a mapsTo of standard → IRI — and encodes back to the same document, attributes included. x-maps-to is the one that pays for itself elsewhere: the mapping engine's autoMatchByMapsTo aligns a field with a differently-named target that shares the IRI, and encoders with a native semantic slot use it directly (LinkML turns it into slot_uri). Registry tooling ignores these attributes; CoreModels reads them. Annotating .avsc files is free interop.

The lossiness inventory

Everything the IR cannot model is declared per occurrence, with a kind, a path, and an explanation. The complete decode-side inventory, with the exact text the coder emits:

TriggerKindExplanation
union with 2+ non-null branchesSemanticNarrowingAvro union of N non-null branches narrowed to the first; alternatives dropped.
mapSemanticNarrowingAvro map has no IR equivalent; approximated as String.
fixedTypeApproximationAvro fixed has no IR equivalent; approximated as String.
logical type decimalTypeApproximationAvro decimal precision/scale not modelled; approximated as Double.
missing or invalid field typeTypeApproximationMissing/invalid Avro field type; approximated as String.

Here they are all at once:

{
  "type": "record",
  "name": "TelemetryEvent",
  "namespace": "com.acme.telemetry",
  "fields": [
    { "name": "id", "type": "string" },
    { "name": "payload", "type": ["null", "string", "bytes"] },
    { "name": "headers", "type": { "type": "map", "values": "string" } },
    { "name": "checksum", "type": { "type": "fixed", "name": "Md5", "size": 16 } },
    { "name": "amount", "type": { "type": "bytes", "logicalType": "decimal",
                                  "precision": 9, "scale": 2 } }
  ]
}

Decoding succeeds — a decode does not hard-fail on a construct it can approximate — and returns four ledger entries: the narrowing on payload, the map on headers, the fixed on checksum, the decimal on amount. Read the fine print they encode. payload still ends up optional: the null branch became optionality before the narrowing chose string over bytes. amount keeps bytes as its preserved token, so re-encoding emits "bytes" — but logicalType: "decimal" and its precision are gone. Approximation is a one-way door: once a map has become String in the IR, re-encoding produces "string", not a reconstructed map. The ledger is the only witness that the door was walked through, which is why every surface returns it.

Round-trip fidelity, stated precisely

For a record built from the constructs in the first table — namespace, doc, x-maps-to, a nullable union, an enum, an array, a timestamp-millis field — decode → encode reproduces the input document, and decode → encode → decode yields a structurally identical IR: same types, elements, value types, requiredness, cardinality. Encoding that second IR again produces byte-identical JSON. The structural stability of that decode → encode → decode loop is the invariant the coder's tests assert; the rest follows from the encoder being deterministic.

Edge cases, pinned down

Non-record roots fail honestly. { "type": "string" }, or a bare enum, is not a record schema: decode fails at path $ with The root of an Avro schema must be a record. An error, not an approximation — there is no sensible Type to build.

Encode needs a root. The encoder looks for _avroRoot, falls back to the schema's first type, and otherwise fails with The schema has no root type to encode as an Avro record. The same message appears in a subtler case: if a mapping renames the root record, the preserved marker points at a type that no longer exists. Keep the record name stable when Avro is the target.

One document, one root. Encoding an IR that holds several unrelated types — a project with two tables, say — emits the root record and everything reachable from it. The others are absent, and no ledger entry announces them. Count your types, or scope the export.

Null defaults are the one asymmetry. Non-null defaults round-trip ("default": 3, "default": "web" come back verbatim). A literal "default": null on a nullable field does not re-emit — the optionality is fully carried by the ["null", T] union, but re-add the attribute if a consumer reads it.

Degenerate unions. A union containing only "null" decodes as an optional String: the least-wrong reading of a field that can hold nothing.

Names travel verbatim. When IR arrives from another format, type and taxonomy identifiers become Avro names untouched. A taxonomy synthesized from a JSON Schema inline enum can carry an id like Customer::status::enum, and that exact string appears as the enum's name. We emit it rather than silently renaming — but Avro's naming rules are stricter than the IR's, so review names before registering such a schema, or rename at the source.

The data side, briefly

The record coder is the schema coder's companion, faithful to Avro's JSON encoding rather than to plain JSON: an optional field is null or a one-key union object ({"string": "SPRING"}), a DateTime is the underlying epoch-milliseconds long, an enum value is its bare symbol, a collection is a JSON array. Two Avro limits are declared rather than smuggled: records carry no identity, so a record id is not written (Avro records carry no id; the record id was not written.), and a document is single-type, so a record of another type is dropped with a StructuralDrop naming both. Fields the schema does not know are declared too.

Why it is built this way

A coarse, honest IR for cross-format reasoning; a per-format extras channel for exact re-emit; declared lossiness wherever the two diverge; semantic annotations as the thread that survives every crossing. Individually small; jointly, the difference between a converter you run and one you trust. The transform documentation covers the companion pieces — the HTTP surface, the mapping engine, and quickstarts for Avro and its sibling formats.