Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -56,6 +56,10 @@ Two folders keep the chain from starting every session at zero. `docs/sessions/`

The wiki ships stocked. Three reference shelves come with the repo: engineering fundamentals (Brooks, Parnas, Naur, and the essential-vs-accidental and wrong-abstraction lenses the challenge rounds cite), design fundamentals (the Norman-to-Rams canon distilled for hierarchy, grouping, type, and color arguments), and motion fundamentals (easing, springs, gesture feel, and motion cost, adapted with credit from Emil Kowalski's and Meng To's MIT-licensed work). Start at [the index](docs/wiki/INDEX.md).

## Optional skills

Foundry also keeps a small set of standalone skills that are useful outside the build chain. [Candidate Assessment](skills/candidate-assessment/SKILL.md) turns rough Product Design interview notes into direct, evidence-led Ashby feedback for hiring manager interviews, portfolio reviews, design exercises, and final decisions. It is deliberately opinionated about early-stage designers: problem finding, a useful spike, customer-led velocity, visible craft judgment, and real AI build practice. Install it by asking Codex to install `skills/candidate-assessment` from this repository.

## What it costs

The full chain is thorough and token-heavy. On a mid-size feature it plausibly lands in the low hundreds of thousands of tokens end to end, which is real money on metered plans. The review rounds are not automatically the expensive part: in the source project's measured runs, build volume, repeated context-heavy reads, duplicate checks, and late mechanical work were often larger contributors. The rounds still found defects late in the chain.
Expand Down
64 changes: 64 additions & 0 deletions skills/candidate-assessment/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
---
name: candidate-assessment
description: Turn rough Product Design interview notes into concise, evidence-led Ashby feedback using Simon Corry's startup hiring bar. Use when Simon asks to run candidate assessment, assess a hiring manager interview, review a portfolio interview, score a design exercise, or synthesize a final hiring decision. Apply only to Product Design candidates Simon is assessing.
---

# Candidate Assessment

Assess Product Design candidates for a small, high-velocity startup. Look for the person who makes the role and the team more ambitious, not the person who most cleanly resembles a conventional designer.

This is a demanding bar. It is not a vibe check and it is not an HR compliance exercise. Strong evidence can be unconventional. Weak evidence does not become strong because the candidate is polished, likable, prestigious, or fluent in design language.

## Inputs

Require:

- Interview type: `hiring manager`, `portfolio review`, `design exercise`, or `final synthesis`
- Target level or role
- Simon's rough notes, transcript excerpts, or supporting paragraphs

Run one interview type at a time. Use `final synthesis` only when evidence from multiple interviews is available.

If the interview type is missing, ask one short question. If the level is missing but obvious from the role title, state the assumption and continue. Do not invent missing observations.

## Route the Assessment

Always read [scoring-and-output.md](references/scoring-and-output.md). Then read only the reference for the requested interview:

- [hiring-manager.md](references/hiring-manager.md)
- [portfolio-review.md](references/portfolio-review.md)
- [design-exercise.md](references/design-exercise.md)
- [final-synthesis.md](references/final-synthesis.md)

Read [research-basis.md](references/research-basis.md) only when revising the rubric or explaining why it works. Do not add research citations to candidate feedback.

Read [ashby-setup.md](references/ashby-setup.md) when creating or changing the actual Ashby interview plans.

## Assessment Method

1. Pull out observed behavior, candidate claims, outcomes, contradictions, and missing evidence.
2. Keep claims separate from proof. Use phrases such as `said they...` when the candidate described something but the interview did not substantiate it.
3. Score each criterion independently against its behavioral anchors and the target level.
4. Apply the core gates for that interview. Do not average scores.
5. Write the Ashby fields in the exact order and format in the scoring reference.

## Simon's Bar

- Prefer a sharp, useful spike over smooth general competence. Name the spike and its likely cost.
- Reward people who find consequential problems and act without waiting for a task list.
- Treat ambiguity as the work. The candidate should create a point of view, not ask the interviewer to finish the brief.
- Pair velocity with learning. Shipping quickly matters when it produces customer evidence and the candidate owns the misses.
- Reward constructive challenge. Agreement, deference, and easy chemistry are not collaboration signals.
- Treat craft as judgment made visible. Look for intent, behavior, hierarchy, detail, and a human pulse. Portfolio polish alone proves very little.
- AI fluency is mandatory. Chat use and one-shot generation do not meet the bar. Look for durable workflows, working builds, model understanding, evaluation, and explicit human judgment.
- Background, pedigree, role labels, accent, charisma, hours worked, and generic culture fit are not evidence.

## Writing Rules

- Sound like a thoughtful design leader recording a decision, not a recruiting department.
- Be candid, specific, and restrained. No inflated praise. No HR euphemisms.
- Use short bullets with a clear reason. Name the evidence and what it means.
- Avoid generic phrases such as `strong communicator`, `great culture fit`, `passionate`, or `impressive background` unless the following words make the claim concrete.
- Do not soften a No into ambiguity. Do not manufacture certainty when evidence is missing.
- Keep the full response between roughly 175 and 300 words unless the evidence genuinely needs less.
- Never mention this skill, AI authorship, or the research in the Ashby-ready output.
4 changes: 4 additions & 0 deletions skills/candidate-assessment/agents/openai.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
interface:
display_name: "Candidate Assessment"
short_description: "Turn interview notes into sharp Ashby feedback"
default_prompt: "Use $candidate-assessment to assess this Product Design interview for Ashby."
64 changes: 64 additions & 0 deletions skills/candidate-assessment/references/ashby-setup.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
# Ashby Setup

Create three Product Design interview plans. Keep the shared fields, replace the generic company traits with the interview-specific criteria below, and use the same 1 to 4 labels everywhere.

## Shared Fields

### Overall Recommendation

- `4`: Strong Yes
- `3`: Yes
- `2`: No
- `1`: Strong No

### Reason for Recommendation

Prompt: `What is the decisive evidence for or against this candidate? Use three to five short bullets.`

### Next Steps

Prompt: `Advance, decline, or name the precise condition for continuing.`

### Areas to Probe

Prompt: `What important question remains unresolved? Maximum three.`

### What Candidate Cares About

Prompt: `What did the candidate explicitly say they want to build, learn, change, or avoid?`

Remove the duplicate general Comments/Notes field from these interview plans. Keep recruiter-only fields outside the design scorecard. Granola Enhanced Notes can remain as raw source material but should not replace interviewer judgment.

## Hiring Manager Scorecard

- `Distinctive Edge and Generative Curiosity`: Do they bring a useful spike, original product perspective, and questions that make the conversation better?
- `Problem Finding, Autonomy, and Impact`: Do they notice consequential problems, create momentum without a task list, and produce a real outcome?
- `Velocity, Learning, and Product Judgment`: Do they ship the smallest useful move, learn from customers, make explicit tradeoffs, and own mistakes?
- `Constructive Challenge, Self-Awareness, and Trust`: Can they disagree plainly, absorb better evidence, know their limits, and preserve trust?
- `AI-Native Builder Practice`: Have they built durable AI-enabled products or workflows, and can they explain model behavior, evaluation, and the human judgment that remains?

## Portfolio Review Scorecard

- `Problem Choice and Framing`: Did they find the real problem beneath the brief and explain why it mattered?
- `Personal Ownership and Judgment`: What did they personally decide, change, and move across the organization?
- `Craft, Taste, and Human Intent`: Does the work show intentional behavior, hierarchy, language, detail, and a human pulse?
- `Shipping, Customer Contact, and Learning`: Did real use change the work, or is this a tidy retrospective story?
- `Distinctive Spike`: What can they do unusually well, and what does that strength cost?
- `AI-Enabled Build Practice`: Did AI materially change how they built, tested, or learned, beyond chat and one-shot generation?

## Design Exercise Scorecard

- `Question Quality and Discovery`: Do their questions expose consequential uncertainty and alter what they do next?
- `Problem Framing and Perspective`: Can they turn a vague prompt into a specific, defensible point of view?
- `Exploration and Distinctive Thinking`: Do they consider meaningfully different directions rather than decorate the first idea?
- `Prioritization, Tradeoffs, and Velocity`: Can they choose, explain what they are sacrificing, and use the hour well?
- `Collaboration and Adaptation`: Can they challenge, listen, and improve the frame when a new constraint arrives?
- `Direction, Craft, and Human Intent`: Does the rough direction cohere, and does it show care for actual human use?

Do not add an AI score to the design exercise unless AI use is explicitly part of that exercise. The final decision still requires AI evidence from another stage.

## Field-Level Scale Description

Use this description on every scored criterion:

`1 Strong No: evidence runs against a core requirement. 2 No: some signal, but below the independence or judgment required. 3 Yes: credible independent performance at level. 4 Strong Yes: rare, repeatable evidence that materially raises the team's bar. Do not average. When something was not observed, leave it unscored and say so in the comment.`
63 changes: 63 additions & 0 deletions skills/candidate-assessment/references/design-exercise.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,63 @@
# Design Exercise

## What This Interview Is For

Watch the candidate create clarity from a deliberately vague prompt. The point is not a finished screen. The point is how they question, frame, choose, explore, and adapt while another person is in the room.

Use a fictional or adjacent problem, not free consulting work. Give a two-sentence prompt with genuine ambiguity. Around the middle of the session, introduce one meaningful constraint that tests the candidate's frame. Answer questions honestly, but do not rescue them with a hidden brief.

The candidate may use paper, Figma, code, or AI after they have made the frame visible. Tool use earns no points on its own.

## Criteria

### Question Quality and Discovery

Core gate.

- `1`: Runs straight to a solution or asks a checklist of questions without using the answers.
- `2`: Covers obvious users and constraints, but does not find the uncertainty that matters.
- `3`: Asks economical questions, turns answers into hypotheses, and deliberately reduces the riskiest uncertainty.
- `4`: Finds the question beneath the prompt and changes what the room believes the exercise is about.

### Problem Framing and Perspective

Core gate.

- `1`: Treats ambiguity as missing instructions and never establishes a point of view.
- `2`: Produces a plausible frame, but it is generic or disconnected from the evidence gathered.
- `3`: Defines a specific user, problem, outcome, and boundary. States assumptions and chooses a defensible perspective.
- `4`: Creates a sharp, surprising frame that opens a materially better product direction.

### Exploration and Distinctive Thinking

- `1`: Fixates on the first familiar pattern.
- `2`: Generates variations, but not meaningfully different approaches.
- `3`: Explores distinct possibilities, compares their consequences, and brings an original idea or useful spike.
- `4`: Reimagines the problem without losing practical contact with the user or business.

### Prioritization, Tradeoffs, and Velocity

- `1`: Cannot choose, or confuses speed with skipping thought.
- `2`: Chooses a direction but cannot explain what was sacrificed or learned.
- `3`: Makes explicit tradeoffs, selects the smallest useful move, and manages the hour well.
- `4`: Uses time and evidence exceptionally well, advancing the idea while keeping the most consequential uncertainty visible.

### Collaboration and Adaptation

- `1`: Defends the artifact, ignores input, or waits for permission.
- `2`: Accepts input but merely complies, or pushes a view without incorporating evidence.
- `3`: Thinks aloud, challenges assumptions respectfully, uses the interviewer as a collaborator, and adapts without losing coherence.
- `4`: The exchange improves both people's thinking. The candidate can absorb a disruptive constraint and produce a stronger frame.

### Direction, Craft, and Human Intent

- `1`: The direction is incoherent or careless about actual use.
- `2`: Understandable but generic, with important behaviors or edge cases ignored.
- `3`: The chosen direction is coherent and intentional. Key behavior, hierarchy, language, and human consequences are considered.
- `4`: Even in rough form, the direction shows exceptional taste and makes the intended experience palpable.

## Interpretation

Problem framing is the primary gate. Strong visual output cannot rescue a candidate who needed the interviewer to define the work. A rough outcome can still pass when the reasoning, perspective, and judgment are strong.

Do not score AI unless the candidate uses it or the interview explicitly tests it. If used, note whether it extended thinking or merely generated surface output. Final synthesis still requires separate evidence that the AI bar is met.
36 changes: 36 additions & 0 deletions skills/candidate-assessment/references/final-synthesis.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,36 @@
# Final Synthesis

## What This Assessment Is For

Make the hiring decision across interviews. Reconcile evidence. Do not average scorecards and do not let the loudest interviewer substitute for a decision.

## Required Evidence

A final Yes or Strong Yes requires credible evidence of all four:

- Problem finding and independent ownership at the target level
- Product craft, taste, and human judgment
- Velocity tied to shipping, customer contact, or a real learning loop
- AI-native build practice at 3 or higher

For staff, lead, or manager roles, also require evidence that the person creates leverage beyond their own output. This can be growing others, improving the build system, changing the quality of decisions, or creating a capability the team did not have.

## Synthesis Method

1. Name the candidate's spike in one plain sentence.
2. Name the cost or risk that comes with it.
3. Separate repeated evidence from a single strong story.
4. Resolve contradictions by asking which interview observed behavior closest to the real work. Do not erase the contradiction.
5. Check all required evidence and level expectations.
6. Decide what the team would be able to do with this person that it cannot do now.

## Overall Anchors

- `1: Strong No`: Evidence contradicts a required capability, ownership is not credible, or behavior presents a serious trust risk.
- `2: No`: Interesting strengths exist, but at least one required capability remains below bar or unproved after the process.
- `3: Yes`: Clears every required capability at level, has a useful spike, and the associated risks are manageable and explicit.
- `4: Strong Yes`: Rare. Clears every requirement, brings a distinctive capability, expands the role or company's ambition, and is someone Simon would reorganize meaningful work to hire.

Do not use team similarity as evidence. A candidate can create productive disruption and still be a Yes. The line is whether their challenge produces better work while preserving trust and follow-through.

If an exceptional candidate describes a role that differs from the opening, say so directly in Next Steps. Shaping a role around a rare person is allowed. Pretending a mismatch does not exist is not.
58 changes: 58 additions & 0 deletions skills/candidate-assessment/references/hiring-manager.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,58 @@
# Hiring Manager Interview

## What This Interview Is For

Decide whether this person should work directly on Simon's team. Test whether broad trust will make them more useful or simply leave them waiting for direction. The conversation can roam, but the scoring cannot become a chemistry score.

## Criteria

### Distinctive Edge and Generative Curiosity

Core gate.

- `1`: Rehearsed, derivative answers. No visible curiosity beyond the role in front of them.
- `2`: Has interests and preferences, but cannot turn them into an original product point of view.
- `3`: Has a clear spike, asks questions that change the conversation, and makes grounded connections others may miss.
- `4`: Brings an unusual, deeply developed capability or perspective that makes the possible role bigger.

Do not confuse eccentricity, confidence, or good banter with a spike.

### Problem Finding, Autonomy, and Impact

Core gate.

- `1`: Waits for assignments or describes ownership that collapses under questioning.
- `2`: Can execute a framed problem but depends on others to choose the problem, remove ambiguity, or create momentum.
- `3`: Repeatedly notices consequential problems, frames them, recruits the right people, and gets useful work into the world.
- `4`: Creates opportunities nobody assigned, changes company direction or capability, and can show the resulting impact.

### Velocity, Learning, and Product Judgment

- `1`: Optimizes for output or perfection without a learning loop. Cannot name tradeoffs or mistakes.
- `2`: Moves quickly in familiar conditions, but customer evidence and decision quality are inconsistent.
- `3`: Chooses the smallest useful move, gets it in front of real people, changes course from evidence, and owns what went wrong.
- `4`: Consistently turns uncertain product bets into fast, high-quality learning without becoming reckless or precious.

### Constructive Challenge, Self-Awareness, and Trust

- `1`: Protects ego, blames others, or uses sharpness as permission to diminish people.
- `2`: Either defers too easily or pushes without listening and follow-through.
- `3`: Can disagree plainly, revise their view, name their limits, and preserve trust while moving the work.
- `4`: Makes the room more ambitious. They challenge the premise, absorb better evidence quickly, and help other people find their own best work.

### AI-Native Builder Practice

Core gate. Use the shared AI anchors.

Ask for concrete artifacts and workflows. Two useful prompts are:

- How has AI changed your actual design and build process? What is different now?
- How has your understanding of models changed what you believe products can become?

A chatbot feature, prompt collection, or enthusiasm for Cursor and Claude is not enough. Look for working systems, agent orchestration, evaluation, customer contact, and a developing mental model.

## Motivation and Anti-Sell

Capture, but do not score, what the candidate wants to build and what their version of a perfect role looks like. Be honest about the environment: little task-list management, high autonomy, frequent shipping, real customer contact, and permission to challenge the premise.

The decisive fit question is not `Do they want this job?` It is `Will this environment release their best work?`
Loading