Skip to content

Pax Updates - #727

Merged
jgarzik merged 34 commits into
mainfrom
updates
Oct 7, 2026
Merged

jgarzik merged 34 commits into
mainfrom
updates

Conversation

@jgarzik

@jgarzik jgarzik commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

No description provided.

jgarzik and others added 30 commits October 6, 2026 05:23
Its only caller is cfg(linux, x86_64), so every other host built it as
dead code and clippy --all-targets warned.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every hard-link site (-rwl, copy mode's link to an earlier copy, and
extraction of a typeflag-1 member) went through create_replacing, which on
EEXIST unlinked the destination and retried. When that name was the link
source itself, the source was destroyed: `pax -rwl tree .` deleted every
file in tree, and a multiply-linked file visited twice lost a name.

link_replacing now compares the (dev, ino) of both names first and leaves
an existing name that already is the source in place. -rwl reports
"Unable to link file to itself", as BSD pax does. The writer skips a
member whose hard-link original is its own name, rather than archiving
"f == f", which extracted by deleting f.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
cpio has no link typeflag: the names of a multiply-linked file are separate
members sharing (c_dev, c_ino). The reader returned each as a plain regular
file and extraction wrote them independently, so a newc set -- which GNU
cpio, bsdcpio and initramfs store with the data on the last name only --
came back as empty files beside one full one, with exit 0.

Extraction now tracks link sets (LinkSets, replacing the unused
ExtractedLinks), links later names to the file already created, and when
the data arrives on a later name creates it exclusively and moves the
earlier names over. -v shows later names as "== first".

The writer stored the real inode number masked to the field (18 bits in
odc, 16 in binary), so unrelated linked files could share a c_ino and be
merged by any reader. Members are now numbered per archive, the names of
one file sharing their first name's number.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
-a located the end of the archive with its own parser, which stepped over
members by the raw ustar size field. It did not know a pax size= record
overrides that field -- which is how members over 8 GiB are written -- or
that some types carry no data, so it took a point inside a member for the
end, wrote there and truncated the rest, with exit 0.

Append now walks the archive with PaxReader, the reader -r and list use,
seeking over member data on a seekable file, and writes where the
end-of-archive indicator begins. The same walk supplies the format and the
-u times; append.rs's three ad-hoc header parsers are gone.

The size rule also gave typeflags 3, 4 and 6 their declared size, so a
FIFO or device header with a nonzero size field swallowed the members
after it. POSIX stores no data for them, and for type 6 says the field
shall be ignored. A pax size= record no longer bypasses the rule.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With -p o or -p e, a chown refused with EPERM was only diagnosed, and the
archived mode was then applied with its set-id bits: a non-root
`pax -r -pe` of a root-owned 06755 member, or `pax -rw -pe` of a setuid
binary, left a set-id file owned by the invoking user. POSIX -p: if the
user and group IDs are not preserved for any reason, pax shall not set
S_ISUID and S_ISGID.

chown_result reports whether the chown took, replacing three copies of
the EPERM handling, and AttrPolicy::mode keeps the bits only when it did.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A pathname read from standard input can contain a NUL -- the common slip
`find -print0 | pax -w` -- and can name no file. ftw built a CString from
it with unwrap() and panicked, after pax had already truncated the -f
archive. Copy mode and the tar and cpio front-ends did the same.

ftw now reports such a path through its error reporter instead. pax
diagnoses the name when reading the list, sets exit status 1 and goes on
with the remaining names, as BSD pax does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The readers took any short read of a header for a clean end of archive,
so an archive cut inside a header block -- an interrupted download --
listed and extracted only the members before the cut and exited 0.
bsdtar and BSD pax report it and exit 1.

read_header now treats no bytes at a header boundary as the end and a
partial header as an error, for ustar, pax, cpio and multivolume reads.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Read mode applied -s to a member's name but not to a hard link's target,
which is another member's name: with -s ',^,P/,' the renamed P/b was
linked to whatever sat at the old name a -- an unrelated file, if one
was there. The target now takes the same -s and --strip-components
rewrite as the member, in read and list mode.

Copy mode built each child's name from its parent's substituted name and
applied -s again, so the substitution compounded per directory level
(out/pre/pre/d/x). It now keeps the source names and substitutes each
once, as write mode does. And in write and copy mode a directory whose
name -s maps to the empty string pruned its whole subtree; POSIX ignores
only that name, so its descendants are now processed under their own.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A failed write to the archive came back as the same Io error as a failed
read of a source file, and only BrokenPipe and StorageFull were treated as
fatal. Past a file-size or quota limit pax went on, reporting the error
against every remaining source file. Archive writes now fail with
PaxError::ArchiveWrite, applied at one wrapper in front of the format
writers, which ends the run with one diagnostic.

BlockedWriter restarted a record from byte 0 after a partial write that
was followed by an error, so a retry repeated the bytes already accepted.
It now resumes where the underlying writer stopped, and retries EINTR.

A source read error after the member header was written left the member
short of its declared size, so every later header landed inside its data
and nothing after it could be read. The member is now zero-padded to its
declared size, as is a file that shrank, and a file that grew is cut to it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
PaxReader replaced the global values wholesale at each g header, so a
later g naming only a comment erased an earlier uname and mtime. POSIX
gives the last record precedence for the keywords it names only; each g
is now merged into the globals in force.

A zero-length record is a POSIX deletion, but the reader parsed it as a
value, and for mtime, atime, uid, gid and size an empty value was a fatal
error: `pax -w -x pax -o mtime:=` wrote an archive pax could not read, and
an x record "uid=" stopped every later member. Deletions are now recorded
and remove the earlier extended and global values (and the uname/gname
header fields).

In read and list mode only -o keyword:=value was applied, loosely parsed
after the fact. keyword=value now acts as a global record at the start of
the archive and keyword:=value as a record ending every extended header,
both parsed by the archive's record parser so a bad value is refused. A
pending x header no longer survives a skipped GNU long-name group.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The tar and cpio front-ends read arguments with env::args(), which panics
on a non-UTF-8 argument, and pax took operands as Strings, so a non-UTF-8
pathname was a usage error. Operands, path-valued option-arguments and
the tar -T / cpio -0 name lists (which went through to_string_lossy and
could select a different file) now keep their bytes. tar and cpio share
one argument walker.

-p was a single value although the synopsis is [-p string]...; repeats
now combine, the last conflicting character winning. -H and -L were
independent flags; the last one given now determines the behavior. And
an option-argument may begin with '-', as in -s -abc-xyz-.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Before the blocked reader ran, a 512-byte gzip peek -- through std's 8 KiB
Stdin buffer -- did a short read of the raw archive, and later reads were
a fixed 10240 bytes. On a tape-like input, where each read returns one
record and loses the rest, that dropped data: a member came back as NULs
with exit 0. Every read now asks for more than the largest record, so a
record is never cut, and the record size is whatever a read returns, as
POSIX requires ("blocking shall be automatically determined on input").
Gzip detection looks ahead through the same reader.

Archives to standard output went through std's line-buffered stdout,
which split each record at its last newline. Archive I/O on stdin and
stdout is now plain read(2)/write(2) on a dup of the descriptor, so each
write is one record of the -b size.

With -z, gzip now sits above the blocked writer, so the compressed stream
is what is blocked, and the gzip trailer is written by an explicit flush
whose error is reported instead of ignored in Drop. The reader decodes
member by member and treats record padding after the last member as the
end. Multivolume sizing counts the trailer and last-record padding.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- -o linkdata had no effect: a pax hard-link member now carries the
  file's size and data when it is given.
- A socket became an empty regular file in ustar and pax, exit 0. It is
  now diagnosed and left out (cpio still stores it as a socket); the
  three typeflag mappings are one function.
- A hard link was recorded before its first name was archived, so an
  unreadable a left "b == a" with no a. It is recorded after the member.
- A pre-1970 mtime wrapped to the year 2514, and negative fractional
  times were mis-signed. Times are signed; pax writes an mtime record
  for any time outside the ustar range, and ustar and cpio refuse one.
- The extended-header name for a top-level member was absolute
  ("/PaxHeaders.N/f"): %d and %f now follow dirname and basename.
- An empty archive pax itself wrote could not be read or appended to.
- -t was ignored in copy mode and did not restore directories' access
  times; both are restored, through a descriptor.
- -X left out the mount-point directory itself; POSIX only stops
  descending below it.
- -rwl with -L or -H linked the symlink rather than its target, and a
  link that could not be made across devices was copied but made the
  exit status 1.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Patterns follow XCU 2.14 with the 2.14.3 rules pax specifies: bracket
  classes, equivalence and collating elements; no '*', '?' or bracket
  match of '/' or a leading '.'; an unterminated '[' matches itself
  instead of being fatal. The matcher is iterative and linear, where
  the old one backtracked exponentially on long '*' runs and could
  overflow the stack. Names and patterns are bytes end to end.
- -n used a pattern up on a directory, dropping the hierarchy below it
  that POSIX says it still selects.
- -u compared after -s/-i renamed the member, and a member -u rejected
  still used up its -n pattern. Selection is now select, -u, take, in
  POSIX order (new modes/select.rs replaces two duplicate selectors).
- -c hid the diagnostic for a pattern that matched nothing.
- A directory member could not replace an existing non-directory.
- GNU base-256 numeric fields aborted the read; they are now parsed,
  signed for mtime.
- -s sliced a str at regex byte offsets and panicked under LC_ALL=C on a
  match inside a multibyte character; substitution works on bytes.
- Owner names were looked up for every member, not only under -p o.
- Member data is seeked over on a regular, uncompressed archive file
  instead of read, and under -n reading stops once every pattern is
  used up.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- tar with an empty -T list read names from standard input instead of
  archiving nothing; cpio -t alone was rejected although -t implies -i.
- find -depth | cpio -idm left a 0700 directory at 0755: a directory
  created only to hold members no longer counts as pre-existing for -u
  or -k, so the member's attributes are applied at the end.
- cpio -p and pax -rw -d applied a read-only directory's mode before its
  contents were copied; copy mode now defers directory attributes the
  way read mode does, deepest first, and also after a fatal error.
- Path components were opened O_RDONLY, so extraction failed below a
  search-only (0300) directory; they are opened O_PATH / O_SEARCH. The
  deferred-directory list is smaller and the parent walk is cached.
- EOF on /dev/tty under -i now ends the run in every mode, as POSIX
  requires, instead of failing one file at a time.
- Archive bytes reached the terminal raw in the final fatal message,
  the -v owner columns and two diagnostics; they go through escape.rs.
- -s parsing: an escaped delimiter keeps its literal meaning on both
  sides, \\ before a delimiter parses, and \0 and \x in the replacement.
- -M wrote a symlink's target as member data and could not read the
  archive back; it uses the ustar header builder.
- The -o tokenizer kept escaped commas in values, decodes printf escapes
  in listopt, keeps trailing blanks and concatenates repeated listopt.
- Copy mode now honors -k and -u for existing directories.
- On a terminal, list output is written line by line, and a write error
  on standard output ends the run with one diagnostic.
- The cross-tool tests skipped or passed whenever the system tool failed;
  they now skip only when it is absent and fail when it fails.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- The stdin name list, tar -T and cpio -0 were read whole before the
  first member was written: memory linear in the list (62-115 MiB for
  1M names) and nothing written until the producer exited, so
  find | pax -w | ssh could not stream. Names are now read as the walk
  asks for them (1.5-1.7 MiB). The first name is still read before -f
  truncates an existing archive.
- Extended-header records were stored with a linear search per keyword
  (40k keywords: 6.2 s) and the global records copied into every member
  (2000 globals x 1000 members: 16.6 s). A member's records are now a
  map over one shared copy of the globals (0.04 s; under 0.01 s). path
  and linkpath records over 64 KiB are rejected, as GNU long names and
  cpio names already were, and a member path is parsed into one buffer
  rather than an allocation per component.
- The hard-link tracker shares cpio's LinkSets but still remembers a file
  for the whole run: a name list can reach a name twice, and forgetting
  the file first stores it again in full and can split the pair.
- Copy mode copies data with copy_file_range on Linux and a buffer of up
  to 128 KiB elsewhere, not 8 KiB (1 GiB: 2.3 s to 1.4 s).
- Copy mode re-walked each member's whole destination chain, O(depth^2):
  the tree keeps a descriptor per level of the last chain walked, within
  half the descriptor limit (depth 1500: 72 s to 17 s).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- tar -T names now select members in -t/-x, as pattern operands; an
  empty list selects none
- an unreadable first name in a list fails before the archive is
  created or truncated
- append refuses a zero-led file that is not all zeros instead of
  overwriting it
- append steps over a lone zero block rather than truncating the
  members after it
- append's archive scan no longer emits read-mode diagnostics or sets
  exit status 1
- tar -r/-u keep the archive's format; -a accepts ustar/pax in either
  direction
- a dangling GNU long-name record ends the archive where it begins
- trailing g headers after a dangling x header are kept on append
- -a with -M is refused
- write and append skip the output archive itself, by (dev,ino)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…gless cpio options

- cpio writer: only multiply-linked files take c_ino numbers reserved
  for their set (counting down from the field maximum); other members
  count up and wrap, so links survive past 65535 members in bin;
  diagnose when the sets alone exhaust the field
- cpio writer/reader: remember a link set for the whole archive, so a
  name repeated beyond its link count keeps the set's c_ino and is
  relinked
- reader: a member sharing (c_dev,c_ino) whose data differs in size,
  mode or mtime from its set's is extracted on its own, not linked over
- reader: when the data-carrying last name of a newc set is not
  extracted (unselected, -u, -k, -s, -i), its data still reaches the
  earlier names via a fresh temp file
- odc/bin: refuse device major/minor wider than 8 bits instead of
  masking
- cpio front-end: -I with -o, -O with -i, -F/-I/-O with -p, and -E
  outside copy-in are usage errors
- tests: feed child stdin from a thread so large name lists cannot
  deadlock

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…er ids, deferred dir identity, fatal sink errors

- -rwl under -H/-L: compare the followed source with the destination,
  so a destination that is the link's target is kept, not unlinked
- chown to uid/gid (uid_t)-1 is not an owner restore: diagnose it and
  drop set-id bits
- deferred directory attributes: record (dev,ino), skip a dir that was
  superseded or swapped instead of failing or stamping the wrong one;
  reopen search-only (O_SEARCH, or O_PATH via /proc/self/fd) when the
  dir is 0300
- -t under -L restores the atime of a directory reached through a link
- an implicitly created directory is claimed by the first member
  naming it, so -k leaves later members of that name alone
- a failed parent walk clears the last-parent cache so a stale
  directory fd is never reused
- EDQUOT and EROFS end read/copy; -O stdout write errors are a new
  fatal StdoutWrite and read mode stops on fatal errors
- use file_id() for (dev,ino) everywhere

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…mtimes, refuses self-overwrite

- create copied directories with source mode|0700 (umask applied) via a
  shared make_dir_at, so a 0700 tree is never world-accessible mid-run
- copy mode replaces an existing non-directory with a directory, as
  read mode does
- -u compares against a directory's mtime from before this run touched
  it, so `find -depth | cpio -pdm` and `pax -rw -u` apply mode and
  mtime (read mode too)
- Linux copy_file_range returning 0 on the first call falls back to
  read/write, so procfs/sysfs/overlayfs files are no longer copied empty
- a directory renamed by -s to '.' has its contents copied into the
  destination
- a destination that is the source itself (same dev/ino) is diagnosed
  and skipped instead of rewritten with links split
- -i blank answer for a directory skips only that name; its children
  are still offered (copy and write mode), matching -s and BSD pax
- a refused directory header is diagnosed and its contents are still
  archived

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Any regexec result other than 0 is now no match. macOS returns
REG_ILLSEQ for text invalid in the locale's encoding, which left the
match offsets unset and was taken as an empty match at offset 0 --
so pax -s inserted the replacement into non-UTF-8 names, and grep/sed/
awk addresses matched such lines.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…recision

- mark every pattern operand that matches a member, so overlapping
  operands (d d/x) no longer report "not found"
- -n: keep the selected directory's hierarchy even when the archive
  lists children first, and when the pattern is "."
- patterns: a trailing '/' means directory only and selects its
  hierarchy; 'a/*' no longer matches a/ itself; a '/' inside brackets
  makes '[' literal (XCU 2.14.3); an absolute name's leading '/' is the
  directory '/', not an empty prefix; brackets decode characters under
  LC_CTYPE
- -c with no patterns selects every member
- tar --exclude: wildcards match a leading '.', and excluding a
  directory excludes its contents when listing and extracting
- cpio -i -r: keep a newer file at the name typed at the prompt, not at
  the archived name (pax -u keeps POSIX order: before -s/-i)
- -u compares modification times to the nanosecond (read and copy mode)
- -s takes bytes; g replaces empty matches as sed does

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…imits

- cap the merged global (g) extended-record set at 64 MiB in total
- writers refuse pax path/linkpath records and cpio names over the
  64 KiB reader limit
- honour typeflag 1 data only in archives using pax headers;
  -o linkdata writes a size= record
- ignore the prefix field of old GNU headers
- never split a ustar name at its trailing slash; directory path
  records keep the slash
- refuse base-256 mode/uid/gid/dev fields wider than 32 bits instead of
  truncating
- accept leading-space octal fields and signed checksums; detection
  shares verify_checksum
- read typeflag 0 with a trailing slash and no data as a directory
- accept an archive that ends inside its second end-of-archive block
- validate -o keyword values in write/append mode before writing
- reject extended records without a trailing newline or with a signed
  length
- compose x headers and GNU L/K records in either order; x path
  supersedes L
- round negative pax times with more than 9 fractional digits down
- diagnose socket members on extract instead of skipping silently

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…read-ahead

- -M: volumes are read as one byte stream by the shared reader; the
  separate header loop is gone (GNU long names, pax x/g, lone zero
  block now handled)
- -M: only the last volume ends with the end-of-archive indicator; a
  missing volume runs the script or prompts, then is an error; a stale
  later volume is not read, and a volume found after the end of the
  archive is reported
- -M: GNU volume labels and 'M' continuation headers are read; a set
  that starts with a continuation, or uses GNU.volume.* records, is
  refused
- -M writer: header checked before a volume change; the old volume is
  closed before the script runs or the prompt appears; the script gets
  /dev/null as stdin when pax reads names from it; volume names keep
  the archive name's bytes; -z with -M is refused
- read: cpio block count is bytes consumed; a seekable stdin is left
  just past the archive and its padding
- read: -b may name records over the write limit; reads are 1 MiB
- -z: gzip CRC and length checked at end of archive; last record padded
  only on a device, so no zeros follow the gzip trailer on a pipe
- archive files are closed explicitly and close errors reported
- -w -i: EOF on /dev/tty still finishes the archive with its trailer

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… listopt edge cases

- a usage error's "Try '<prog> --help'." line is written unescaped, so
  on a terminal it stays its own line instead of following a '?'
- -o keyword ends at the first '='; a ':' before it selects the
  per-file form, so comment=a:=b has the value "a:=b"
- -o option-arguments are bytes: values (path:=, listopt, name
  templates) need not be UTF-8; exthdr/globexthdr templates expand as
  bytes
- listopt decodes printf \ddd octal escapes and always appends a
  newline per file (POSIX)
- tar/cpio '--opt=' with an empty value is the empty value, not the
  next argument (tar -c --file= a.txt truncated a.txt)
- tar -c with no operands and no -T is refused, as GNU tar does;
  tar -r/-u with none append nothing; neither reads stdin
- -i reply is read as bytes and only its newline is stripped
- -u compares member names by their bytes; lossy keys made distinct
  non-UTF-8 names collide. Drop the now-unused rawpath::MatchName

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ect their bugs

- a non-ASCII path or linkpath now always gets a pax record (ustar
  header fields are the portable character set); only a non-UTF-8 name
  is declared hdrcharset=BINARY. The non-ASCII path-record test no
  longer passes only because of a sub-second mtime
- linkdata test checks each hard link carries a data block, not just a
  size
- link-onto-itself copy test asserts exit status and diagnostics, and
  uses file operands so it reaches the linkat EEXIST path
- test ustar fixtures encode numeric fields in base-256 when octal
  overflows, so a uid >= 2097152 no longer panics
- unmatched bracket-pattern test asserts exit 1 and the diagnostic
- path-record limit test uses a name of exactly 64 KiB and one byte
  over
- repeated -p test checks that both strings apply
- GNU base-256 test actually encodes the size field in base-256

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… empty -T

- -i opens /dev/tty before write/append create or truncate the archive (or
  the first -M volume), so a run without a terminal leaves it intact.
- The files the archive is written to are a shared set; -M adds each volume
  as it is created, so the walk no longer stores a volume as a member.
- An empty tar -T list selects every member and the archive is always
  opened, as bsdtar and GNU tar do; a missing archive is an error again.
- The self-link check treats ./h/a and h/a as one name.
- Append -u compares to the nanosecond when the member records a fraction.
- tar -c skips a socket with a warning and exits 0; pax mode still errors.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ords, refuse wide cpio owners

- The "regular, trailing slash, size 0 = directory" rule is applied to the
  final path and size, after extended-header records, not the raw name and
  size fields; the writer no longer cuts a non-directory's name field at a
  slash.
- size is always written with its real value: -o delete= and size:= can no
  longer strip it, and path/linkpath survive delete when the ustar fields
  cannot hold them.
- A leading '+' is rejected in pax numeric records and cpio hex fields.
- odc and bin refuse uid, gid (and bin nlink) values too wide for the field
  instead of masking them; newc checks uid/gid too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…data, chmod without following

- Only archive, -O stdout and terminal failures end a run; ENOSPC, EDQUOT,
  EROFS and EPIPE on a destination file are per-member errors that name the
  member, as POSIX CONSEQUENCES OF ERRORS requires.
- -n keeps reading until every started newc link set has its data, so the
  first name is no longer left empty.
- A link set created empty, with no earlier name left to link to, waits for
  its data on a later name instead of linking every name to the empty file.
- Extracted and copied files are closed explicitly and a close failure is
  reported against the member.
- FIFO and device modes are set via one chmod_at helper that never follows
  a symlink (AT_SYMLINK_NOFOLLOW; O_PATH + /proc/self/fd on Linux).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…a followed source link

- A copy-mode directory whose destination is itself is no longer a reason to
  skip its subtree when -s or -i may rename its children: it is left as is
  and the walk descends. Without -s/-i one diagnostic still covers it.
- Under -H/-L a destination that is the followed symlink itself counts as
  the source, for files, directories and -l links, so `pax -rw -H link .`
  no longer unlinks the link and puts an empty directory in its place.
- One followed_link helper replaces three copies.
- Remove copy mode's unreachable pattern selection and matches_any.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
jgarzik and others added 2 commits October 6, 2026 19:46
…cters after empty matches

- plib's regex takes an explicit not-BOL flag; pax -s with g passes it once
  a match has been made, so ^ no longer re-matches the rewritten start.
- sed steps past an empty match by one whole character (shared plib locale
  helper) instead of one byte, so later matches on a UTF-8 line are no
  longer lost to REG_ILLSEQ.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
No '*' stands for the empty filename below '/', so '/*' names what is in
'/', not '/' itself, whose hierarchy took in /.hidden past the
leading-period rule.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@jgarzik
jgarzik requested a balanced review from Copilot October 7, 2026 00:41
@jgarzik jgarzik self-assigned this Oct 7, 2026
Comment thread pax/tests/common/mod.rs Fixed

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The change set is large and spans subtle, high-impact areas (regex anchoring/encoding semantics, archive write error classification, byte-level CLI parsing, and compression framing) that warrant final human review despite no concrete defect being found.

Review effort: Balanced
Findings: None

What changed in this PR

This PR ("Pax Updates") pairs a substantive rewrite of sed's substitution engine with the supporting plib/ftw primitives it needs (and a broader set of pax refactors referenced by the title). The visible core change replaces sed's large, fragile range-tracking substitution machinery (~440 lines of match_pattern/filter_groups/update_pattern_space helpers) with a single streaming substitute loop that walks matches left to right, handles empty matches and multibyte stepping under LC_CTYPE, and builds replacement text in append_replacement. plib::regex gains a matched() helper (so a regexec error such as macOS REG_ILLSEQ is treated as no match, not an empty match at offset 0) and a captures_at_bytes_notbol variant that controls ^ re-anchoring via REG_NOTBOL, which is what makes the new global-substitute loop correct.

Changes:

  • Rewrote sed's s/// substitution into a streaming substitute/append_replacement pair, correcting empty-match semantics, g/Nth counting, ^/$ anchoring, and multibyte stepping.
  • Added plib::regex::matched() and captures_at_bytes_notbol(); removed the local REG_NOMATCH constant in favor of the result == 0 check.
  • Added/used plib::locale::next_char_offset for whole-character advancement, with new sed integration tests covering global empty matches, multibyte locales, anchors, and replacement groups.
File Description
text/​sed.rs Replaces range-tracking substitution with streaming substitute/append_replacement; simplifies execute_replace and address matching.
text/​tests/​sed/​mod.rs Adds locale-aware substitution tests and corrects a prior expectation (whole match replaced, not just the group).
plib/​src/​regex.rs Adds matched() and captures_at_bytes_notbol() for explicit BOL anchoring; drops the redundant REG_NOMATCH constant.
plib/​src/​locale.rs Provides next_char_offset for multibyte-safe stepping over empty matches (used by sed).

(The PR title also covers a large pax refactor — error-handling semantics, ArchiveSink error relabeling, CLI OsString byte-handling, GzipReader multi-member handling, member selection, and ftw NUL-in-path reporting — which I reviewed for consistency and found internally coherent and well-tested.)


💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

jgarzik and others added 2 commits October 6, 2026 21:06
When the empty first name of a set was replaced by an unrelated member,
the set's file was unlinked and ext4 handed its inode number to the next
file created, which the (dev, ino) check then took for the set's file:
on Linux the replacing member ended up with the set's data. Holding the
file open while the data is pending keeps its inode number its own.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Clears a CodeQL cleartext-logging alert on the test helper.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@jgarzik
jgarzik merged commit 22702eb into main Oct 7, 2026
20 checks passed
@jgarzik
jgarzik deleted the updates branch October 7, 2026 02:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants