Skip to content

CBOR: preservation-aware encode reuses stale chunk sizes, writing different content #575

Description

@solidsnakedev

Summary

When a byte or text string was decoded as indefinite-length chunks, the captured format keeps the chunk sizes. If the value then changes and is encoded with that format, the encoder writes the old chunk sizes over the new value. The output is valid CBOR carrying different content, and no error is raised:

const { format } = CBOR.fromCBORHexWithFormat("5f4201024103ff") // h'010203' as chunks 2 + 1

CBOR.toCBORHexWithFormat(new Uint8Array([0xaa]), format)
// 5f42aa004100ff, decodes as h'aa0000' (zero padded)

CBOR.toCBORHexWithFormat(new Uint8Array([1, 2, 3, 4, 5]), format)
// 5f4201024103ff, decodes as h'010203' (tail dropped)

Text strings behave the same way: "x" under the format of "hij" decodes as "x\0\0", and a change can split a UTF-8 character, which gives undecodable output.

This breaks the preservation spec, .specs/cbor-encoding-preservation.md:

  • L243: only a compatible format branch may be reused
  • L244: preserved widths that no longer fit must fall back

Integer widths already follow L244 and strings do not. The comment at Transaction.ts L242 also says reconciliation "handles structural changes gracefully".

Affected

packages/evolution/src/CBOR.ts

  • encodeBytesSync indefinite branch (L1156-1176): chunk lengths come from stringEncoding.chunks, never from the value, and the zero-filled buffer is sized from them
  • text string indefinite branch (L1247-1265): same pattern

Reachable through CBOR.toCBORHexWithFormat and every *WithFormat encoder built on it, such as Transaction.toCBORHexWithFormat. The automatic cache path (toCBORHex after fromCBORHex, addVKeyWitnesses) does not change existing strings, so it does not reach this.

Fix

Reuse a captured chunk layout only when it still fits the value:

  • bytes: the chunk sizes add up to the value's length
  • text: the same, and every chunk ends on a UTF-8 character boundary

Otherwise fall back to the default encoding for that string, as integers already do. The definite branch keeps its header width, since encodeIntHeader already falls back when the length does not fit.

Regression test

  • given: format captured from 5f4201024103ff
  • before fix: encoding h'aa' gives 5f42aa004100ff, and h'0102030405' gives 5f4201024103ff
  • after fix: both decode back to the value that was encoded
  • text: "x", "hijk" and "héllo" under the format of 7f626869616aff decode back to themselves
  • control: an unchanged value under its own format stays byte-identical

Devnet, raw submission through Ogmios, using the documented edit flow:

  1. decode a transaction whose metadata is {0: h'010203'} in chunks, with Transaction.fromCBORHexWithFormat
  2. set the metadata to h'aa' and the body's auxiliary data hash to AuxiliaryData.toHash of the new metadata
  3. encode with Transaction.toCBORHexWithFormat and the original format

The SDK writes the auxiliary data as a1005f42aa004100ff. The node rejects it with 3107, metadata hash mismatch: provided df360c50… (the SDK's hash of h'aa'), computed 9c397db0… (the hash of the bytes written, which hold h'aa0000'). The original transaction is accepted.

Must FAIL on main today and PASS after the fix.

Reference

Blocks the fix for #576 (Plutus data encodings), which will pass the captured format into the bounded bytes encoder.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions