Hire Calibrate

Run the pre-loop calibration session — 60 minutes with ≥3 raters, review the rubric, score 2 anchor candidates the team has previ…

frankxai 6 updated 22d ago
Claude CodeGeneric
View source ↗
---
name: hire-calibrate
description: Run the pre-loop calibration session — 60 minutes with ≥3 raters, review the rubric, score 2 anchor candidates the team has previously seen, surface inter-rater drift, agree on hire-bar examples. The calibration session itself is the deliverable. Without it, the loop runs on uncalibrated rulers. Per Project Oxygen — cross-rater alignment beats rater quality. Not legal advice.
allowed-tools: Read, Write, Grep, Glob
argument-hint: role-slug (required, ICP and interview architecture must exist) + --rater-count <3|4|5|6+> + optional context paragraph on team's prior calibration history
---

# /hire-calibrate

This is part of the People Intelligence reference vertical. Composes with Genius Profile + Vision/Brand for company-as-candidate framing.

Load `SIP.md`, `VOICES.md`, `agents/starlight-hiring.md`, `skills/people-intelligence/structured-hiring.md`, the existing ICP (`people-intelligence/hiring/icp-<role-slug>-*.md`) and interview architecture (`people-intelligence/hiring/interview-<role-slug>-*.md`). Produce a **Calibration Session Script** — a facilitator-ready 60-minute agenda. Hand off to running the actual loop.

## Disclaimer (non-waivable)

**Hiring decisions touch employment law and protected-class considerations. This is system architecture, not legal advice. The calibration session itself does not surface legal questions; question stems must already have been reviewed by counsel before this session runs.**

This command produces the script. The facilitator runs the live session. The team enters the loop calibrated, not running on instinct.

## Input
$ARGUMENTS

## Flags

- `--rater-count <3|4|5|6+>` — number of raters in the calibration session. Minimum 3 (per Project Oxygen, cross-rater alignment requires ≥3 reference points). 4-5 is the sweet spot. 6+ creates discussion drag.

## Process

1. **Disclaim.** Open with the non-waivable disclaimer.

2. **Locate.** Confirm role-slug. Read the ICP and interview architecture. If either is missing, halt and route to upstream command.

3. **Identify anchor candidates.** Two candidates the rater team has previously interviewed for similar roles — one retrospective hire-yes (worked out), one retrospective hire-no (or worked out poorly). The calibration session scores these against the new rubric. Their actual outcomes are the calibration anchor.

4. **Build the 60-minute agenda.**
   - **0:00 - 0:05 (5 min) — Frame the session.** Why calibration matters. Project Oxygen finding: cross-rater alignment beats rater quality. Decision: this team will not run an uncalibrated loop.
   - **0:05 - 0:20 (15 min) — Rubric walk-through.** Each rater paraphrases what 1, 3, and 5 mean for each dimension. Surface mismatched paraphrases. Agree on language.
   - **0:20 - 0:40 (20 min) — Anchor candidate scoring.** Each rater independently scores the two anchor candidates against the new rubric (8 min independent). Then surface scores in plenary (12 min). Where did raters disagree by ≥2 points on any dimension? Discuss those specifically.
   - **0:40 - 0:55 (15 min) — Hire-bar agreement.** What does a 3-on-this-dimension look like in our actual team? What does a 5-on-this-dimension look like? Agree on examples for each load-bearing dimension.
   - **0:55 - 1:00 (5 min) — Question stem commitment.** Each rater commits to using the agreed first-question stems verbatim. This kills divergent question framing across raters.

5. **Surface bias-pattern primer.** Before the loop runs, name the bias patterns the facilitator will flag in the debrief: halo, similarity-attraction, first-impression, recency, contrast effect, confirmation bias. Raters do not have to self-diagnose; the facilitator names patterns out loud.

6. **Decision-rule pre-commitment.** Pre-commit:
   - Hire-or-no-hire decision rule (median score ≥4 on ≥75% of dimensions, ≥3 on all dimensions, or whatever this team agrees to)
   - Tie-breaks default to hire-no
   - Inter-rater dispersion ≥2 points on a load-bearing dimension triggers re-interview, not vibes-resolution

7. **Save.** Write `people-intelligence/hiring/calibration-<role-slug>-<YYYY-MM-DD>.md`.

8. **Hand off.** Run the loop. After the loop, run `/hire-debrief <candidate>`.

## Output format

```markdown
# Calibration Session Script — <Role Title> — <YYYY-MM-DD>

> **Hiring decisions touch employment law and protected-class considerations. Question stems must already have been reviewed by qualified counsel before this session runs.**

## Context

- **Role:** <title>
- **ICP file:** `people-intelligence/hiring/icp-<role-slug>-<date>.md`
- **Interview architecture file:** `people-intelligence/hiring/interview-<role-slug>-<date>.md`
- **Rater count:** <3 | 4 | 5 | 6+>
- **Facilitator:** <name — typically the hiring manager or an HR partner>
- **Anchor candidates selected:** <yes — names redacted in this artifact / no — flagged>

## Why this session

Per Project Oxygen and the multi-rater research literature: cross-rater alignment matters more than individual rater quality. A team with mediocre raters but tight calibration produces better hire signal than a team with star raters running on individual instinct. We are not running this loop until we are calibrated.

## 60-minute agenda

### 0:00 - 0:05 — Frame the session (5 min)

**Facilitator opens:**

> "We are not running an uncalibrated loop. Calibration is the difference between a rubric and a vibe scale wearing numbers. The next 55 minutes are the highest-leverage 55 minutes of this entire hire — they prevent the drift that produces miss-hires. We will:
> 1. Walk the rubric and surface where we paraphrase it differently.
> 2. Score two candidates we've all seen against the new rubric.
> 3. Discuss where we disagree.
> 4. Pre-commit to question stems and the decision rule.
>
> Three rules: structured scores before discussion always. Tie-breaks default to hire-no. The facilitator names bias patterns in real time so no one has to self-diagnose."

### 0

Maintain Hire Calibrate?

Let people know it's listed here — add the badge (live metrics, light/dark aware) or a plain link to your README or docs.

[Hire Calibrate on getagentictools](https://getagentictools.com/loops/frankxai-hire-calibrate?ref=badge)