Deterministic proto3 Pipelines: The Plan Is the Artifact
A conversion you run once is a favor. A conversion you run on every commit is infrastructure, and infrastructure needs two properties that one-off tooling never has to prove: the second run must produce the same bytes as the first, and when the input changes in a way the conversion cannot honor, the pipeline must stop rather than quietly produce something different.
Deterministic proto3 Pipelines: The Plan Is the Artifact
A conversion you run once is a favor. A conversion you run on every commit is infrastructure, and infrastructure needs two properties that one-off tooling never has to prove: the second run must produce the same bytes as the first, and when the input changes in a way the conversion cannot honor, the pipeline must stop rather than quietly produce something different.
The CoreModels transform surface is built on both. This article shows how to turn proto3 conversions
into pipeline steps — deterministic outputs, plans as committed artifacts, drift that fails the
build, and a batch pattern for a directory of .proto files.
Property one: the same input produces the same bytes
The proto3 encoder writes its output by hand in a fixed order — syntax, package, imports, file
options, then messages and enums, with each message emitting its options, reserved statements,
fields, nested enums, and nested messages in that sequence. Nothing depends on dictionary iteration
order or on the wall clock. Decode a file and encode it twice and you get identical text; encode the
result of decoding your own output and you get the same text again. It is a fixed point, not merely
a repeatable function.
That property is what makes generated artifacts safe to commit. Check the produced JSON Schema or
.proto into the repository next to the source, regenerate in CI, and git diff --exit-code becomes
a meaningful gate: a non-empty diff means a real change to the schema, never churn from the
generator.
One nuance for .proto sources specifically. A comment-free file that already follows the
encoder's canonical layout survives a proto3 round trip byte-for-byte, including package, imports,
options, reserved statements, field numbers, optional labels, oneof groupings, and enum
numbers. Comments and bespoke formatting do not come back — the tokenizer consumes them and they
are not part of the schema — so if your baseline file carries a license header or its own layout,
compare the regenerated file against a regenerated baseline rather than against the hand-written
original.
Property two: the plan is a first-class artifact
Every mapping call returns the plan it executed. For an identity conversion of a shipment event schema with seven fields, that is nine operations — one type mapping, one element mapping per field, one taxonomy mapping — each stamped with where it came from. Four of them:
{
"operations": [
{ "kind": "TypeMapping", "origin": "Inferred",
"sourceTypeId": "ShipmentEvent", "targetTypeId": "ShipmentEvent",
"targetLabel": "ShipmentEvent" },
{ "kind": "ElementMapping", "origin": "Inferred",
"sourceElementIds": ["ShipmentEventShipmentId"],
"targetElementIds": ["ShipmentEventShipmentId"] },
{ "kind": "ElementMapping", "origin": "Inferred",
"sourceElementIds": ["ShipmentEventSignature"],
"targetElementIds": ["ShipmentEventSignature"] },
{ "kind": "TaxonomyMapping", "origin": "Inferred",
"sourceTaxonomyId": "ShipmentStatus", "targetTaxonomyId": "ShipmentStatus",
"targetTreatment": "InlineEnum" }
]
}
"origin": "Inferred" means a heuristic matched it by label and compatible type. That is a fact
worth having in code review: the mapping is not hidden inside a script, it is a document you can
read, diff, and approve.
Save it and replay it. plan/execute takes the same source, the plan as a string, and a target
format:
jq -n --rawfile p shipment_event.proto --rawfile q plans/shipment_event.plan.json '{
sourceFormat: "protobuf",
sourceSchema: $p,
plan: $q,
targetFormat: "protobuf"
}' > replay.json
curl -sS -X POST \
"https://coremodels.example.com/graph/transform/plan/execute/$PROJECT_ID" \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
--data-binary @replay.json > replayed.json
A stored plan earns no shortcut: it passes the identical validation gate as a freshly produced one before anything executes. Replay the same plan against the same source as many times as you like and the output is byte-identical every time.
Editing the plan: a curated projection
Because the target is exactly what the plan projects, editing the plan is how you author a projection — no script, no template, no second copy of the schema.
Take the identity plan above and delete the one operation that maps ShipmentEventSignature. Replay
it. The emitted .proto is the original minus that field:
syntax = "proto3";
package acme.logistics;
import "google/protobuf/timestamp.proto";
message ShipmentEvent {
string shipment_id = 1;
uint64 sequence = 2;
ShipmentStatus status = 3;
optional string carrier_note = 4;
repeated string package_ids = 5;
google.protobuf.Timestamp occurred_at = 6;
}
enum ShipmentStatus {
SHIPMENT_STATUS_UNSPECIFIED = 0;
PICKED_UP = 1;
IN_TRANSIT = 2;
DELIVERED = 3;
}
The surviving fields keep their original numbers — 1 through 6 — because those numbers were preserved from the source, and a removed field's number is simply not re-used. And the omission is not silent. The ledger gains:
{ "kind": "StructuralDrop",
"path": "Type[ShipmentEvent].ShipmentEventSignature",
"explanation": "Element is a member of the mapped type but no operation maps or drops it." }
That message is the engine's rule stated out loud: a field that belongs to a mapped type but has no operation is reported, never quietly omitted. If you meant to drop it, the record is your receipt. If you did not, it is your bug report.
The same rule cuts the other way, and it is the one thing to know before authoring a plan by hand:
only what the plan maps reaches the target. A plan with element mappings but no TypeMapping
produces a valid .proto with a package and no messages, because no message was ever mapped. When
you write operations yourself, map the type first.
Drift that fails the build
Here is the payoff that ad-hoc conversion scripts cannot give you.
Store a plan. Six weeks later somebody renames a field in the source .proto. Your pipeline replays
the stored plan, and the gate refuses it before the engine runs:
{
"success": false,
"lossiness": [ /* the source-decode ledger still appears here */ ],
"errors": [
{ "path": "plan.Operations[1](ElementMapping).SourceElementIds",
"message": "Source element 'ShipmentEventNoSuchField' does not resolve." }
],
"schema": null
}
The path names the operation index and the operation kind; the message names the construct that
vanished. A script would have produced an artifact missing a field and exited zero. This exits with
success: false and tells you which field and which operation, which is a five-minute fix instead of
a downstream incident.
The gate also rejects a plan that produces the same target id twice, a plan that both maps and drops the same source node, a named field transform that is not registered, and an element mapping whose arity or types do not line up. Every strategy passes through it — inferred, explicit, AI-proposed, and stored alike.
The three mapping kinds, applied to proto3
inferred(the default) matches source constructs to a target hint by label and compatible type. It suits proto3 unusually well, because a field's label is its wire name:shipment_idmatchesshipment_id.caseInsensitivedefaults totrue, which is what you want when snake_case proto fields meet a camelCase or upper-case target. Every operation it emits is stampedInferred.explicittakes a SIA mapping guide:autoMatchByMapsTo,fieldMappings({sourceElementIds, targetElementIds, transformName}),taxonomyDirectives, anddrops. It is the tool for the cases inference cannot express — joining two fields into one, splitting one into two, forcing a taxonomy treatment, deleting something deliberately. A multi-fieldfieldMappingsentry must name a transform from the registry; a plain one-to-one entry may omit it. Its automatic backbone ismapsTo, the shared cross-standard URI two constructs declare; a bare.protofile carries none, so pair an explicit guide with a hint whose constructs are annotated, or author the type alignment as plan operations.aiasks a server-side proposer for a plan and then puts it through the same gate, with at most one repair attempt. It requires Editor or Admin membership and a server-configured key, and it sends the schema content to an external model API — a deliberate choice for a proprietary.proto, and one the tooling advertises rather than hides.
Whichever kind produced it, what you store and replay afterwards is the plan. That is the point of having one plan format.
Batch conversion, end to end
A directory of .proto files, one call each, artifacts and ledger collected:
#!/usr/bin/env bash
set -euo pipefail
HOST="https://coremodels.example.com"
mkdir -p build plans
: > build/lossiness.tsv
for f in schemas/*.proto; do
name=$(basename "$f" .proto)
jq -n --rawfile p "$f" '{
sourceFormat: "protobuf",
sourceSchema: $p,
targetFormat: "jsonschema",
targetHintFormat: "protobuf",
targetHintSchema: $p,
mapping: { kind: "inferred" }
}' > "build/$name.request.json"
curl -sS -X POST "$HOST/graph/transform/schema/map/$PROJECT_ID" \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
--data-binary @"build/$name.request.json" > "build/$name.result.json"
jq -e '.success' "build/$name.result.json" > /dev/null
jq '.schema' "build/$name.result.json" > "build/$name.schema.json"
jq '.plan' "build/$name.result.json" > "plans/$name.plan.json"
jq -r --arg n "$name" \
'.lossiness[] | "\($n)\t\(.kind)\t\(.path)\t\(.explanation)"' \
"build/$name.result.json" >> build/lossiness.tsv
done
Then gate on the ledger rather than on exit codes alone. Different teams draw the line in different places; a strict one fails on anything structural and lets type approximations through with a warning:
if cut -f2 build/lossiness.tsv | grep -qx StructuralDrop; then
echo "A structural drop appeared in a conversion — review build/lossiness.tsv" >&2
exit 1
fi
cut -f2 build/lossiness.tsv | sort | uniq -c
For a directory of proto3 files, that count is usually dominated by two entries you should expect
and can whitelist: unresolved import clauses, and TypeApproximation records for uint32,
uint64, sint32, sint64, the fixed-width encodings, and bytes. Those are the wire facts no
other format carries. What you are watching for is the record that appears for the first time — a
new oneof flattened into optional fields, a map approximated as a string, a service block
dropped, or a StructuralDrop naming a field your plan forgot.
Two closing notes on pipeline hygiene. Plans are source-specific: element ids derive from message and field names, so a plan belongs to the file it was produced from — store it beside that file, not in a shared directory. And a plan's replay is only as stable as the source it names, which is precisely why the gate checks before the engine runs.
The transform section of the CoreModels documentation has the endpoint list, the role each route needs, and the full mapping-guide vocabulary.