Apache Avro logoQuickstart

Your First Avro Transform: One Call, Three Things to Read

You have an `.avsc` file. Somewhere downstream, somebody needs the same shape as JSON Schema, or as a table, or as a data contract — and they need to know what the conversion cost. This article gets you from that file to a converted schema in a single HTTP call, and then teaches you to read the three parts of the answer: the schema, the plan, and the lossiness ledger.

Your First Avro Transform: One Call, Three Things to Read

You have an .avsc file. Somewhere downstream, somebody needs the same shape as JSON Schema, or as a table, or as a data contract — and they need to know what the conversion cost. This article gets you from that file to a converted schema in a single HTTP call, and then teaches you to read the three parts of the answer: the schema, the plan, and the lossiness ledger.

Nothing here writes to a project. The endpoint we use is stateless: it decodes, maps, encodes, and hands everything back. You need a host (https://coremodels.example.com throughout), a bearer token in $TOKEN, a CoreModels project id in $PROJECT_ID (32 hex characters — it scopes authorization only, and Viewer access is enough), plus curl and jq.

Step 1 — the schema

Save this as SensorReading.avsc. It is a small, ordinary record that happens to exercise the five Avro constructs that behave interestingly in any conversion: a namespace, a doc string, an enum, a logical type, and a nullable union.

{
  "type": "record",
  "name": "SensorReading",
  "namespace": "com.acme.telemetry",
  "doc": "One reading from a field sensor.",
  "fields": [
    { "name": "deviceId", "type": "string", "doc": "Stable hardware id." },
    { "name": "reading", "type": "double" },
    { "name": "unit", "type": { "type": "enum", "name": "Unit",
                                "symbols": ["celsius", "fahrenheit"] } },
    { "name": "recordedAt", "type": { "type": "long", "logicalType": "timestamp-millis" } },
    { "name": "note", "type": ["null", "string"], "default": null }
  ]
}

The format key for this file is avro, and it works in both directions — CoreModels decodes Avro and encodes Avro, so the same key is valid as a source and as a target.

Step 2 — the rule, then the call

One rule first, because it is the most common stumble. The default mapping strategy, inferred, matches your source against a target hint and declines rather than guesses when there is none. Call without a hint and you get success: false with the error The inference resolver requires a target IR to match against.

For a straight format conversion the hint you want is the source itself. That produces an identity plan — every construct maps to its own counterpart — and lets the target coder do the format work.

The schema travels inside JSON as a string, so build the body with jq --rawfile rather than escaping quotes by hand:

jq -n --rawfile s SensorReading.avsc '{
  sourceFormat: "avro",
  sourceSchema: $s,
  targetFormat: "jsonschema",
  targetHintFormat: "avro",
  targetHintSchema: $s,
  mapping: { kind: "inferred" }
}' > request.json

curl -sS -X POST \
  "https://coremodels.example.com/graph/transform/schema/map/$PROJECT_ID" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  --data-binary @request.json > result.json

That is the whole transform. The response is one envelope with success, lossiness, errors, the produced schema, and the executed plan.

Step 3 — read the schema

jq '.schema' result.json. For JSON-shaped targets the field holds a JSON object, ready to write to a file:

{
  "type": "object",
  "properties": {
    "deviceId": { "type": "string" },
    "reading": { "type": "number" },
    "unit": { "enum": ["celsius", "fahrenheit"] },
    "recordedAt": { "type": "string" },
    "note": { "type": "string" }
  },
  "required": ["deviceId", "reading", "unit", "recordedAt"]
}

Read it against the input and every decision is visible. The record became the object type and each field a property. The double became number. The enum became an inline controlled list. The timestamp-millis logical type became a date-time value, which JSON Schema carries as a string. And the ["null", "string"] union on note became optionality — which is why note is the one field missing from required. In Avro, a plain field always carries a value; only a union with null means "may be absent", and that is exactly the distinction the conversion preserves.

Step 4 — read the plan

jq '.plan.operations[0]' result.json:

{
  "kind": "TypeMapping",
  "origin": "Inferred",
  "sourceTypeId": "SensorReading",
  "targetTypeId": "SensorReading",
  "targetLabel": "SensorReading"
}

The plan is the list of operations that actually executed, in order. For this call there are seven — one type mapping, five element mappings, and a taxonomy mapping for Unit with "targetTreatment": "InlineEnum". Each is stamped with an origin, and Inferred means "matched by label and type" — a heuristic, surfaced rather than hidden.

Keep this object. POST /graph/transform/plan/execute/{projectId} replays a stored plan against the same source and produces the same output, which is what turns a one-off conversion into a pipeline step you can trust.

Step 5 — read the ledger

jq '.lossiness' result.json returns [] here. That empty array is a real result, not a placeholder: every construct in SensorReading had an exact home on the other side.

Now change one word — "targetFormat": "sql", adding "vendor": "postgres" — and run the same call again. The schema comes back as DDL, in the schema field as a string:

CREATE TABLE "SensorReading" (
  "deviceId" VARCHAR(255) NOT NULL,
  "reading" NUMERIC NOT NULL,
  "unit" VARCHAR(255) NOT NULL,
  "recordedAt" TIMESTAMP NOT NULL,
  "note" VARCHAR(255)
);

and this time the ledger is not empty:

[
  { "kind": "ConstraintRelaxation",
    "path": "Element[SensorReading.unit]",
    "explanation": "Postgres has no inline enum; emitted as VARCHAR (the allowed-value constraint is not enforced)." }
]

One entry, with a kind, the exact path, and a plain-English explanation. The unit column exists; the guarantee that it only ever holds celsius or fahrenheit does not. If that matters, you add a check constraint downstream — and now you know to.

This is the habit worth forming on day one: success: true means "it ran", never "nothing changed." The four kinds you will meet are StructuralDrop (no home in the target), TypeApproximation (a close-but-not-exact type), ConstraintRelaxation (a rule that could not be enforced), and SemanticNarrowing (meaning narrowed or guessed). Treat the list as a review checklist, not an error report.

Step 6 — and back again

Because avro is also a target format, the same call with "targetFormat": "avro" returns your record:

{
  "type": "record",
  "name": "SensorReading",
  "namespace": "com.acme.telemetry",
  "doc": "One reading from a field sensor.",
  "fields": [
    { "name": "deviceId", "type": "string", "doc": "Stable hardware id." },
    { "name": "reading", "type": "double" },
    { "name": "unit", "type": { "type": "enum", "name": "Unit",
                                "symbols": ["celsius", "fahrenheit"] } },
    { "name": "recordedAt", "type": { "type": "long", "logicalType": "timestamp-millis" } },
    { "name": "note", "type": ["null", "string"] }
  ]
}

Namespace, doc strings, the enum with its symbols in order, the exact logical type, the nullable union — all back where they started, with an empty ledger. One nuance is visible if you look closely: the explicit "default": null on note is not re-emitted. The field's optionality is fully carried by its ["null", "string"] union, but if a consumer of yours reads that attribute, re-add it.

The full target list is jsonschema | shex | avro | jsonld | sql | osi | osi-json | owl | linkml | protobuf | odcs | synapse, and only targetFormat changes between them.

Where to go next

Two continuations. If the schema should become a governed model rather than a converted file, POST /graph/transform/schema/import/{projectId} with {"format": "avro", "schema": "<the .avsc text>"} writes it into a project as a Type with its Elements and Taxonomies — that route writes, so it needs Admin. And if you would rather have an AI agent do this, the same engine is an MCP tool called transform_schema, taking the same argument names — with one difference: the nested mapping object is flattened to mappingKind, guide, and caseInsensitive.

The transform section of the CoreModels documentation has the complete endpoint list, the role for each, and ready-to-paste bodies for every format.