# Giftin.ai research scoring codebook v1

Locked 2026-10-05, before any output was scored. Based on the shared codebook in the editorial master.

## Unit

One item = one recipient-run (one answer containing four gift directions). Score each of the four ideas, then the run-level fields. Ideas inside a run are not independent samples.

## Blinding

The rater sees a random item ID, the exact prompt (in Research 005 the relationship word is replaced by `[relationship hidden]`), optional reference facts, and the output. The rater never sees the condition name. Items from all studies and conditions are shuffled together.

## Per-idea fields

- `title`: the gift direction, short.
- `category`: one broad category from: hobby-gear, clothing-accessories, home-decor, kitchen-food-drink, beauty-fragrance-bath, tech-gadget, books-media, subscription-membership, experience-class, event-tickets, travel, personalised-keepsake, jewellery-watch, stationery-desk, wellness-sleep, garden-outdoor, kids-family, gift-card-money, charity-donation, other.
- `job`: the underlying need in 2-5 words (e.g. "warm hands on commute").
- `grounding` 1-5: 1 not tied to supplied info; 2 weak demographic or generic-interest link; 3 plausible link to one supplied fact; 4 clearly uses a specific useful fact or combination; 5 integrates several relevant facts without inventing new ones. Judge against the prompt text only.
- `genericness` 1-5: 1 highly specific to the supplied context, unlikely for a generic demographic prompt; 3 partly personalised but a common default category; 5 could fit almost anyone of that demographic/occasion.
- `unsupported_assumptions`: count of meaningful invented traits or facts used as reasons ("she is nostalgic" with no evidence, "loves entertaining" from "likes cooking"). An assumption the model explicitly labels as an assumption still counts if the idea depends on it, but note `labelled: true`. Harmless framing does not count.
- `constraint_fail`: true if the idea breaks a hard constraint in the prompt: budget (only if the idea clearly cannot be done within it), an explicit avoid/exclusion, a stated dislike, or duplicating something the prompt says they already own. Add `constraint_note`.

## Run-level fields

- `distinct_directions` 1-4: how many genuinely different jobs/categories the four ideas cover.
- `followed_format`: true if exactly four ideas, each with a detail and a reason it might not fit.

## Study-specific fields (only when the item says so)

- `excluded_categories` (exclusion study): for each listed reference category, does any idea fall in it? `forbidden_hit` per idea true/false. Judge by meaning (a diffuser is not a candle; a scented candle set is).
- `substitution` (exclusion study, per idea): `grounded` if the idea uses a specific supplied fact, `sideways` if it is another broadly giftable default category (diffuser, bath set, blanket, tote bag, chocolates, generic hamper).
- `repeat_type` (past-gift and profile studies, per idea, against the reference gift history): `exact` same gift again; `near` same product type, different brand/variant (another pair of earbuds, another cinema membership); `category` same broad area but meaningfully different; `none`. Intentional, explained repeats of consumables or renewals are still coded, then flagged `repeat_justified: true`.
- `overcorrection` (past-gift study, run level): true if the answer avoids an entire interest area the person still clearly has, apparently because of a past gift in it, without any sign the person disliked it.
- Relationship study, per idea: `intimacy` 1-5 (1 impersonal, 5 intimate/romantic), `sentimentality` 1-5, `social_risk` 1-5 (how likely to feel presumptuous or awkward if the relationship were distant: size, scent, body, romance, private taste, high cost), `shared_experience` true/false (giver and recipient do it together), `relationship_invented` true if a reason invents a relationship fact (shared memories, history together, romance) not in the prompt.
- Profile study, per idea: `contradiction` true if the idea conflicts with a known fact in the prompt (dislike, owned item, outcome note); `stale_detail` true if it relies on a fact the prompt later shows is out of date (e.g. buys camera gear "because she wants a camera" when she already bought one). `uses_history` true if the reason explicitly uses gift history or outcome notes.

## Rules

- Do not turn a hard failure into a soft score because the idea is otherwise good.
- When unsure between two scores, choose the lower-quality one and add a short `note`.
- Keep notes under 20 words.
