Skip to content

feat(skill): K-1 code-mode domain globals and K-2 skill auto-arming - #62

Merged
its-janghoon merged 3 commits into
developfrom
feature/k1-k2-code-mode-globals-and-skill-arming
Oct 2, 2026
Merged

its-janghoon merged 3 commits into
developfrom
feature/k1-k2-code-mode-globals-and-skill-arming

Conversation

@its-janghoon

@its-janghoon its-janghoon commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

What

Phase 6 items K-1 and K-2 from the delivery loop.

  • K-1 (50ba34c723, already on the branch) — typed domain objects as code-mode globals.
  • K-2 (0dfe0fbe0b) — skill auto-arming from frontmatter keywords and URL globs.

K-1: typed domain objects, not a tool list

A skill document is cheap only because the objects it drives already exist. Code Mode could
previously only be written as tools.<namespace>.<tool>(...), so a skill had to describe a
tool call rather than name a method. Now a skill body reads:

const body = await page.text()
await channel.send({ text: body.slice(0, 200) })

The mechanism, in @redrob-code/codemode: globals on ExecuteOptions names top-level
tool namespaces that are also bound as bare identifiers in the interpreter's global scope.
A global is an alias for the same tool path, so there is one implementation and one
authorization point, not two. Code Mode stays host-neutral — it never learns what a given
name means. Builtin global seeding moved into seedBuiltinGlobals() and
BUILTIN_GLOBAL_NAMES is derived from it, so a builtin added later cannot silently become
available as a host global name. assertValidGlobals refuses a name that is not a namespace
of the tool tree, one that shadows a builtin, a duplicate, a non-identifier, and a tool
rather than a namespace.

The objects, in packages/redrob/src/tool/domain.ts — real TypeScript interfaces with a
single implementation each:

export kind
Page interface — url / text / query / click / type / navigate
Channel interface — id / send
PageNode interface — one node returned by page.query
DomainUnavailableError tagged error carrying object + missing
unavailablePage the only Page implementation (see below)
sessionChannel the working Channel implementation
unavailableChannel Channel for the catalog preview, where nothing may be posted
pageTools / channelTools / domainTools the namespaces handed to Code Mode
DOMAIN_GLOBALS ["channel", "page"] — the names bound as bare globals

Injected in packages/redrob/src/tool/code-mode.ts via
tools: { …, ...domainTools({ page, channel }) } with globals: DOMAIN_GLOBALS. The domain
namespaces are spread after the MCP catalog deliberately: a connected MCP server named
page must not shadow the object a skill is written against.

page is NOT implemented, and says so

Nothing in the engine process controls a browser page today. There is no page-control seam
from the browser to the engine, so there is nothing for a real Page to call.

Its single implementation is therefore unavailablePage, whose every method fails with
DomainUnavailableError naming the missing capability:

The page domain object is not available in this session: the engine process has no
browser-page control surface; the browser must expose page control to the engine first

It returns no plausible value — not an empty string, not an empty node list — so a skill
written against the interface fails loudly instead of silently reading nothing. The
interface, the tool schemas and the generated instructions all exist and are exercised now,
so skills can be written and reviewed before the browser side lands; only the act fails.

channel is implemented: channel.send appends a text part to the assistant message
that owns the execution, so every surface the session is attached to renders it without the
engine knowing which surface that is. channel.id returns the session id.

unavailableChannel is a second deliberate refusal rather than a gap: it backs the catalog
preview path, so the generated instructions list the same globals a live execution binds
without a preview being able to post into a conversation.

K-2: frontmatter

Two optional keys beside the existing name / description / slash:

---
name: google-docs
icon: https://example.test/icon.png          # a URL
autoInject:
  keywords: [summarise, draft]               # matched against the prompt
  url: [docs.google.com/document/**]         # globs
---

Both lists inside autoInject are optional, and an empty list matches nothing — so
autoInject: {} is a skill that still never arms, rather than one that arms on
everything. A skill that arms on everything is a skill that is always in the prompt.

K-2: the arming decision

A pure function, in packages/core/src/skill/arming.ts:

export const arm = (input: Input): ReadonlyArray<Armed>
export const armedNames = (input: Input): ReadonlyArray<string>

type Input = {
  readonly skills: ReadonlyArray<Armable>   // { name, autoInject? }
  readonly prompt: string
  readonly url?: string | undefined
}
type Armed = {
  readonly name: string
  readonly keywords: ReadonlyArray<string>  // the keywords that matched
  readonly urls: ReadonlyArray<string>      // the globs that matched
}

It reads nothing, caches nothing and logs nothing, so the decision is identical in the
prompt builder, in a UI preview and in a test. arm returns the patterns that armed each
skill, which is what a "why is this skill loaded?" surface needs; armedNames is for
callers that only want the set.

Matching rules, both case-insensitive:

  • keyword — matches on word boundaries, so ai does not match inside said, while a
    keyword like .pdf or docs.google.com still matches in running text.
  • url — * and ? stop at a /; ** crosses path segments; every other character is
    literal, so a . in a hostname is a dot and not a wildcard. The scheme is ignored on both
    sides, so globs are written docs.google.com/document/**, not with an https:// prefix.

A skill with no autoInject block never auto-arms. It stays explicitly loadable.

K-2: one document at a time, and the reason is kept

Loading decodes each skill document on its own, so a malformed frontmatter costs that one
document and not the whole set.

The decoders are Schema.decodeUnknownResult, not decodeUnknownOption: the Option form
answers only "no", which makes a malformed document unfindable among the ones that loaded.
Every rejection is logged at warning level with the file and the decoder's own reason:

SkillV2.load dropped 1 of 3 skill documents from directory:/… (loaded 2, ignoredAutoInject 1):
  /…/bad-type.md: frontmatter rejected: Expected boolean, actual "yes please"

A malformed autoInject is narrower still — it is ignored with its own warning and the
skill loads without it, which leaves the skill exactly where a skill with no block already
sits, instead of losing the skill to a typo in an optional block.

The reasons go in the log message, not in annotations. The default logger renders an
annotation object as [object Object], which was dropping precisely the detail that makes
a drop findable — caught by the test that asserts on captured log output.

Tests

packages/core/test/skill/arming.test.ts — 13 pass, 0 fail, 27 expect() calls.

Covering: a keyword hit; which keyword armed; case-insensitivity in both directions; a
keyword not matching inside a longer word; a URL-glob hit; scheme-ignoring and
case-insensitive URL matching; three glob near-misses that must not arm (documents vs
document, a single * refusing to cross /, a literal dot); a skill with no block;
an empty block; no current URL; keyword and URL hits on one skill; result ordering; and a
malformed autoInject that is ignored rather than fatal while a sibling undecodable
document is dropped with a logged reason and count.

Reverse-verified — each of five restored defects was confirmed to fail the suite:

restored defect result
explicit-only guard removed (no-block skills arm) caught, 5 fail
single * crosses / caught, 1 fail
glob compiled as a regex (dot is a wildcard) caught, 1 fail
malformed autoInject made fatal caught, 1 fail
reason text dropped from the drop log caught, 1 fail

Gates

  • bun turbo typecheck — 18/18 packages successful (also run by the pre-push hook).
  • bun run lint (oxlint) — 0 errors.
  • packages/core: bun test — 987 pass, 0 fail, 119 files.
  • packages/schema: 13 pass, 2 fail in test/event-manifest.test.ts. Pre-existing and
    unrelated: measured with git stash push -u, which gives an identical 13 pass / 2 fail on
    the clean tree. Not touched here.

Not in this change

  • page cannot act. Its interface and its refusal ship; the capability does not, because the
    browser-to-engine page-control seam does not exist yet. Detailed above.
  • K-1's row in the loop names page, screen, mail, vault. Only page and channel are
    here, which is the scope the item was given: the two nothing else depends on. screen,
    mail and vault are not stubbed, not declared, and not silently half-present.
  • No consumer is wired to arm yet — K-2 is the mechanism and the frontmatter, and the
    call site in the session prompt builder is a separate item.
  • icon is decoded and carried on Skill.Info for surfaces that list skills; no surface
    renders it yet.

Follow-up commit: the generated client artifact

CI's core (linux) job was red and it was a real defect in this branch, not flake. The job
runs the core suite (which passed) and then bun run check:generated in packages/client,
which is bun run generate && git diff --exit-code -- src/generated src/generated-effect.
packages/client commits its generated SDK, so adding icon and autoInject to
SkillV2.Info left the committed artifact stale and that gate failed. unit (linux) is only
the aggregation job reporting the same failure, not a second one.

Fixed in feat(client): K-2 regenerate the client types for the new skill keys by running the
repo's own generator rather than editing the file, since the gate rejects a hand edit by
construction. The entire diff is the propagation of the schema change onto SkillsListOutput:

     readonly slash?: boolean
+    readonly icon?: string
+    readonly autoInject?: { readonly keywords?: ReadonlyArray<string>; readonly url?: ReadonlyArray<string> }
     readonly location: string

check:generated exits 0 after the commit.

Gate measurements re-run on the final tree

Every number below was measured on this branch's tip, not carried over:

gate result
bun turbo typecheck --force (codemode, core, redrob, schema) 4/4 successful
whole-repo bun turbo typecheck (pre-push hook) 18/18 successful
packages/codemode bun test 276 pass, 0 fail
packages/core bun test 987 pass, 0 fail
packages/redrob bun test 3399 pass, 0 fail, 22 skip, 1 todo
packages/client bun test 15 pass, 1 fail (pre-existing, see below)
packages/schema bun test 13 pass, 2 fail (pre-existing, see below)

New test files, measured individually:

file result
packages/core/test/skill/arming.test.ts 13 pass, 0 fail
packages/codemode/test/globals.test.ts 13 pass, 0 fail
packages/redrob/test/tool/domain.test.ts 11 pass, 0 fail

Both pre-existing failures were measured against a control rather than inferred, by checking
out origin/develop into a detached worktree and running the same suite there:

  • packages/schema test/event-manifest.test.ts — 13 pass / 2 fail on this branch and
    13 pass / 2 fail on clean origin/develop. Identical.
  • packages/client "exposes every standard HTTP API group" — fails on clean
    origin/develop as well. The control tree in fact shows 14 pass / 2 fail, one failure
    more than this branch, so this change does not add a client failure.

A skill document is cheap only because the objects it drives already exist.
Code Mode could previously only be written as `tools.<namespace>.<tool>(...)`,
so a skill had to describe a tool call instead of naming a method.

Mechanism, in @redrob-code/codemode: `globals` on ExecuteOptions names top-level
tool namespaces that are ALSO bound as bare identifiers in the interpreter's
global scope. A global is an alias for the same tool path, so there is one
implementation and one authorization point, not two. Code Mode stays host-neutral:
it never learns what a given name means. The builtin global seeding moves into
seedBuiltinGlobals(), and BUILTIN_GLOBAL_NAMES is derived from it, so a builtin
added later cannot silently become available as a host global name.
assertValidGlobals refuses a name that is not a namespace of the tool tree, one
that shadows a builtin, a duplicate, a non-identifier, and a tool rather than a
namespace. The generated instructions gain a "Domain globals" section stating the
bare and tools forms are the same call.

Objects, in packages/redrob/src/tool/domain.ts: Page and Channel as real
TypeScript interfaces with a single implementation each, injected under the names
`page` and `channel`.

- Channel is implemented: `channel.send` appends a text part to the assistant
  message that owns the execution, so every surface the session is attached to
  renders it without the engine knowing which surface that is. `channel.id`
  returns the session id.
- Page is NOT implemented, and says so. Nothing in the engine process controls a
  browser page today, so its single implementation is unavailablePage, whose every
  method fails with the capability named: "the engine process has no browser-page
  control surface". It returns no plausible value, so a skill written against the
  interface fails loudly rather than reading an empty string.

The domain namespaces are spread after the MCP catalog deliberately: a connected
MCP server named `page` must not shadow the object a skill is written against.
Two optional frontmatter keys beside name/description/slash:

  icon: string                                      # a URL
  autoInject:
    keywords: [string]                              # matched against the prompt
    url: [string]                                   # globs, e.g. docs.google.com/document/**

Both lists are optional and an empty list matches nothing, so `autoInject: {}`
is a skill that still never arms rather than one that arms on everything.

The arming decision is a pure function in packages/core/src/skill/arming.ts:

  SkillArming.arm(input: Input): ReadonlyArray<Armed>
  SkillArming.armedNames(input: Input): ReadonlyArray<string>

It takes the loaded skills, the current prompt text and the current tab URL,
and returns which skills arm, with the patterns that armed each one. It reads
nothing and caches nothing, so the decision is identical in the prompt builder,
in a UI preview and in a test. Keyword and URL matching are both
case-insensitive; keywords match on word boundaries so `ai` does not match
inside `said`; `*` and `?` stop at a `/` while `**` crosses segments; a `.` in
a glob is a literal dot and not a wildcard. The scheme is ignored on both
sides, so globs are written without one.

A skill with NO autoInject block never auto-arms. It stays explicitly loadable,
which is the point: a skill that arms on everything is always in the prompt.

Loading decodes one document at a time, so a malformed frontmatter costs that
one document and not the set. The decoders are the Result-returning form rather
than the Option form, because the Option form answers only "no" and makes the
cause unfindable; every rejection is logged at warning level with the file and
the decoder's own reason. A malformed autoInject is narrower still: it is
ignored with its own warning and the skill loads without it, since that leaves
the skill exactly where a skill with no block already sits. The reasons are in
the log MESSAGE rather than in annotations, because the default logger renders
an annotation object as [object Object] and loses precisely that detail.

Tests in packages/core/test/skill/arming.test.ts: 13 pass, 0 fail. They cover
a keyword hit, a URL-glob hit, three glob near-misses that must not arm, a
skill with no block, an empty block, case-insensitivity in both directions, and
a malformed autoInject that is ignored rather than fatal while a sibling bad
document is dropped with a logged reason. Each of five restored defects
(explicit-only guard removed, `*` crossing `/`, glob compiled as a regex,
malformed autoInject made fatal, reason text dropped from the log) was confirmed
to fail the suite.

Note: packages/schema has 2 failing tests in test/event-manifest.test.ts. They
fail identically on a clean tree (git stash push -u, 13 pass / 2 fail both
ways) and are unrelated to this change.
`packages/client` commits its generated SDK and gates it with
`check:generated` (bun run generate && git diff --exit-code), so adding
`icon` and `autoInject` to `SkillV2.Info` in packages/schema left the
committed artifact behind and that gate red on CI - the core (linux) job
passed its tests and then failed on this step, and unit (linux) is only
the aggregation job reporting it.

Regenerated with the repo's own generator (`bun run generate` in
packages/client), not hand-edited, because the gate rejects a hand edit by
construction. The whole diff is the two new optional fields on
SkillsListOutput, which is exactly the propagation of the schema change:

    readonly icon?: string
    readonly autoInject?: { readonly keywords?: ReadonlyArray<string>
                            readonly url?: ReadonlyArray<string> }

Measured after the change: packages/client typecheck passes; its suite is
15 pass / 1 fail, and that one failure ("exposes every standard HTTP API
group") also fails on a clean origin/develop worktree, so it is not from
this change.
@its-janghoon
its-janghoon merged commit b7bc394 into develop Oct 2, 2026
19 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant