ODCS logoMCP

"Turn Our Data Contract Into an Avro Schema": ODCS Through `transform_schema`

"The platform team publishes the orders contract — can you make the Avro schema for the streaming team?" That sentence, said to an agent with no tools, produces plausible-looking `.avsc` with invented decisions baked in. Said to an agent connected to CoreModels over MCP, it becomes a call to `transform_schema`: a deterministic engine does the conversion, and the agent gets back the schema, the executed plan, and a machine-readable account of what the trip cost. This article is the complete loop for ODCS — connection, the exact tool contract, two real conversions, and the contract-specific habits that separate a good agent from a confident one.

"Turn Our Data Contract Into an Avro Schema": ODCS Through transform_schema

"The platform team publishes the orders contract — can you make the Avro schema for the streaming team?" That sentence, said to an agent with no tools, produces plausible-looking .avsc with invented decisions baked in. Said to an agent connected to CoreModels over MCP, it becomes a call to transform_schema: a deterministic engine does the conversion, and the agent gets back the schema, the executed plan, and a machine-readable account of what the trip cost. This article is the complete loop for ODCS — connection, the exact tool contract, two real conversions, and the contract-specific habits that separate a good agent from a confident one.

Connection

CoreModels serves MCP at /mcp over streamable HTTP with OAuth 2.0. Clients discover the authorization server from protected-resource metadata and register dynamically — no pre-shared client id; spec-compliant clients run the PKCE flow themselves.

claude mcp add --transport http coremodels https://coremodels.example.com/mcp

In Claude Desktop or claude.ai, it is Settings → Connectors → Add custom connector with the same URL; a JSON-configured client needs only:

{
  "mcpServers": {
    "coremodels": { "type": "http", "url": "https://coremodels.example.com/mcp" }
  }
}

transform_schema is stateless and Viewer-level, so it lives on the public endpoint. The project id it takes scopes authorization only — nothing is read from or written to the project.

The tool contract

The input schema declares additionalProperties: false: a misspelled argument is rejected, not silently ignored.

ArgumentRequiredMeaning
graphProjectIdyesproject id, pattern ^[a-f0-9]{32}$
sourceFormatyesjsonschema | shex | avro | jsonld | sql | osi | osi-json | owl | linkml | protobuf | odcs | odm
sourceSchemayesthe source text — for ODCS, the contract YAML (or JSON)
targetFormatyesjsonschema | shex | avro | jsonld | sql | osi | osi-json | owl | linkml | protobuf | odcs | synapse
targetHintFormat / targetHintSchemanoformat and text of the schema to map toward
mappingKindnoinferred (default) | explicit | ai
guidenoexplicit: the SIA mapping-guide JSON; ai: optional free-text guidance
caseInsensitivenoinferred matching ignores label case (default true)
vendornosql output only
synapseOrg / synapseName / synapseVersionnosynapse output only

Three ODCS-specific facts an agent should hold. First, odcs is valid in both sourceFormat and targetFormat — contracts go in and come out. Second, there are no ODCS-specific arguments: nothing on this tool sets a contract's id, version, or status. When ODCS is the target, the head comes from facts preserved off an ODCS source, or from deterministic defaults (v3.1.0 / 1.0.0 / active) — so a generated contract head is a draft to edit, and the agent should say so. Third, for a straight conversion the idiom is to pass the source as its own targetHintSchema, because inferred mapping declines rather than guesses without a hint.

First ask: contract to Avro

The user pastes the orders contract — head, servers block, six properties, a quality check on order_total (the contract from our ODCS quickstart). The agent's call:

{
  "graphProjectId": "3f2a9c1e5b7d48a0b6c2e4f8091a3d57",
  "sourceFormat": "odcs",
  "sourceSchema": "<the contract YAML, as a string>",
  "targetFormat": "avro",
  "targetHintFormat": "odcs",
  "targetHintSchema": "<the same YAML>",
  "mappingKind": "inferred"
}

The tool answers with one JSON document — success, schema, plan, lossiness. Avro is JSON-shaped, so schema is an object:

{
  "type": "record",
  "name": "orders",
  "fields": [
    { "name": "order_id", "type": "string" },
    { "name": "customer_ref", "type": "string" },
    { "name": "order_total", "type": "double" },
    { "name": "placed_at", "type": { "type": "long", "logicalType": "timestamp-millis" } },
    { "name": "item_count", "type": ["null", "long"] },
    { "name": "gift", "type": ["null", "boolean"] }
  ]
}

The mapping is legible line by line: required contract properties became plain Avro fields (a plain field always carries a value), the two optional ones became ["null", T] unions, number became double, integer became long, and the timestamp logical type crossed to timestamp-millis. The ledger has two entries, both from the reading side: the servers block and the quality check were preserved verbatim in the transform's ODCS channel and declared as SemanticNarrowing — an .avsc has nowhere to put either. The agent's report should carry exactly that: "Here is the record schema. The contract's server bindings and its nullValues mustBe 0 check on order_total do not exist in Avro — enforce that check in the pipeline that consumes the topic." An agent that hands over the schema without the second sentence has discarded the reason the contract existed.

Second ask: the same contract as proto

"Also a .proto for the edge service, please." One argument changes — targetFormat: "protobuf":

syntax = "proto3";

message orders {
  string order_id = 1;
  string customer_ref = 2;
  double order_total = 3;
  string placed_at = 4;
  optional int64 item_count = 5;
  optional bool gift = 6;
}

This time the ledger has three entries: the same two from decoding the contract, plus one from writing the proto — TypeApproximation at Element[ordersPlacedAt]: "proto3 has no date-time scalar; DateTime degrades to string." Same source, different target, different cost, and the ledger is per-call, so the agent never has to remember which target loses what.

Third ask: the tool in reverse

A different user, the opposite direction: "We have a JSON Schema for Member — draft me an ODCS contract so we can start the governance conversation."

{
  "graphProjectId": "3f2a9c1e5b7d48a0b6c2e4f8091a3d57",
  "sourceFormat": "jsonschema",
  "sourceSchema": "{ \"$id\": \"Member\", \"type\": \"object\", \"title\": \"Member\", \"properties\": { \"member_id\": { \"type\": \"string\" }, \"joined_on\": { \"type\": \"string\", \"format\": \"date-time\" }, \"points\": { \"type\": \"integer\" }, \"active\": { \"type\": \"boolean\" } }, \"required\": [\"member_id\", \"joined_on\"] }",
  "targetFormat": "odcs",
  "targetHintFormat": "jsonschema",
  "targetHintSchema": "<the same JSON>",
  "mappingKind": "inferred"
}

The schema value is a YAML string — ODCS is a text format — and it is a spec-valid v3.1.0 contract:

apiVersion: v3.1.0
kind: DataContract
id: sia-schema
name: Member
version: 1.0.0
status: active
schema:
  - name: Member
    properties:
      - name: member_id
        logicalType: string
        required: true
      - name: joined_on
        logicalType: timestamp
        required: true
      - name: points
        logicalType: integer
      - name: active
        logicalType: boolean

The body is exactly right — types crossed, required crossed, format: date-time became timestamp — and the head is exactly a default: minted so the output satisfies the spec's required fields, not because anyone decided this contract is version: 1.0.0 and active. The agent's move here is to present the YAML and flag the head for editing before anyone registers it. That is not a workaround; it is where contract identity decisions belong.

Habits for agents holding contracts

Relay the ledger in domain terms. For most formats the ledger is about types. For ODCS it is about governance: servers, quality checks, SLAs, team, relationships — the promise around the schema. When those appear as SemanticNarrowing on decode, the honest summary is "the schema crossed; the contract's obligations did not — here is where each one now needs a home."

Round-trip when the target is ODCS, and only then. Preserved contract facts (head, physical types, quality, unknown sections) re-emit exactly when the target is odcs. They do not leak into any other format. An agent asked "will we lose anything?" can answer precisely: nothing on an ODCS-to-ODCS trip; the declared entries otherwise.

Escalate mappingKind deliberately. inferred is default and side-effect-free. explicit takes an authored guide (autoMatchByMapsTo, fieldMappings, taxonomyDirectives, drops) — right when the user can state the mapping. ai asks a server-side Claude proposer for a plan and carries real preconditions: a server-configured Anthropic key, Editor or Admin membership on the scoping project, and the schema content travels to the Anthropic API server-side — the tool advertises itself as an open-world interaction for exactly that reason. The proposal passes the same validation gate as every other plan, with at most one repair attempt. State all of this before pointing ai at somebody's proprietary contract.

Keep the plan. The response's plan replays via the HTTP plan/execute route for identical output — the difference between an answer in a chat and a step in a pipeline.

One placement note to close: there is no separate project-export MCP tool for ODCS — when an agent needs a contract out of format-shaped inputs, transform_schema with targetFormat: "odcs" is the route, and the HTTP schema/export endpoint covers exporting a project's governed model as a contract. The MCP quickstart in the CoreModels documentation has connection details and the full tool inventory.