EditionKR · 한국어EN · EnglishJA · 日本語ID · Bahasa IndonesiaTH · ไทยZH · 简体中文ZH · 繁體中文
📖 You're reading the full book for free. If it helps, the polished e-book is a way to support the author → Amazon Kindle $9.99  ·  Leanpub — 30% off  ·  Author on LinkedIn
GAME DESIGN × AI WORKFLOW AI Workflow for Game Designers No Fabricated Numbers A Six-Month Field Manual — Claude Code · Prompts · Validation · Production Memory 304 atoms 48 skills 85 chapters Minsoo Lee DESIGN DIRECTOR · 24Y · 2026

AI Workflow for Game Designers

No Fabricated Numbers — A Six-Month Field Manual for Claude Code, Prompts, Validation, and Production Memory

The 304 rules, 48 tools, and automation systems that a director with 24 years in the industry actually ran on a mid-sized (10–50 person) team — exactly as they were used

By Minsoo Lee · 2026

Everything in this book is current as of the first half of 2026. Pricing, models, and features of AI tools change quickly, so please check each official page for the latest figures and installation steps.


Distribution and Copyright

This book is released free of charge, in the hope that it will be read widely. But whether you are reading the Korean original or an English or Japanese translation, one fact must stay attached to it everywhere: the original author of this book is Minsoo Lee (이민수).

You are free to do the following. Personal study, noncommercial sharing and quotation, noncommercial translation, and use in internal study groups — provided you credit the original author, include a link to the canonical edition, and state clearly if you have modified the content or carried it into another language. Derivative works must be shared under the same terms.

Please ask for permission first. Commercial publication or sale, use as course material in paid classes, inclusion in a company's products or services, and any redistribution that removes the author credit.

The knowledge in this book is complete in this one volume, and it is free as it stands. If it helps you and you would like to support the author, the official e-book and the paid kits are a way to do that.


Preface — A Book Written Three Times

This book was written three times.

The first version was not a book at all — it was an internal company manual. Over six months of running an AI workflow at work, I pinned decisions, tools, and procedures down in writing so the team would never have to ask about the same rule twice. It was not made for publication; it was an operations document, built up to cut down on daily repetition. That is where this book's concreteness comes from — it is not based on examples invented for a book, but on a manual that was actually in service.

The second was the first manuscript that turned that manual into a book. It ended up as a plausible-sounding survey of "look what AI can do." It had a lot of tables, and a lot of numbers demonstrating impact. Most of those numbers carried a note in small print: "illustrative figures." When I reread it, that was its biggest flaw. It talked about using AI without ever showing a real screen, and it talked about impact while citing made-up numbers. In moving the manual into book form, I had managed to lose the very concreteness that made it worth writing.

So I wrote the whole thing again, a third time. That is the text you are holding now. I brought back the concreteness of the internal manual, and I kept the book's principles simple.

First, every chapter shows you a real session from start to finish. The full prompt I typed, the raw output the AI produced, what I rejected in that output, and how I made it redo the work — all of it is here. No chapter ends with the sentence "AI takes care of it."

Second, every number is one of three kinds. A public standard anyone can verify (model token pricing, accessibility guidelines), a constant actually present in my system's code, or a value explicitly labeled "this is my estimate." There is not a single fabricated savings table. I made honesty the differentiator.

Third, I quote, as is, the real system I ran in production for six months. The 304 decision cards (atoms), the 48 tools (skills), the hook that automatically pulls in relevant memory on every input, the memory structure that lets one person run four people's worth of collaboration context — all of it. Not an abstract "some tool," but the actual file names, code, and scores, written down as they are. In the worked examples I masked only company and project names and the names of teammates; the concreteness of the workflow is untouched. (The company that gave permission to publish this book is named openly in the acknowledgments — because they agreed to it.)

I am a game designer with 24 years in the industry. I got my start doing QA and review on single-player games, and went on to direct an MMORPG that ran in dozens of countries, to work on the early development of a 200-person AAA MMORPG, and to handle live ops for a global mobile MMORPG — a career spent building RPGs, MMORPGs, and their variants. On projects like Ragnarok Online, Bless Online, and the MIR series, I served as director, design lead, and systems designer, and at times as PM; I also once founded a small mobile game company.

To be honest, I was made a director too early. After that came a long stretch back in the trenches: working data sheets and combat numbers hands-on as a systems designer, producing quests and NPCs line by line as a content designer, planning events, churning out new content. Many of the workflows in this book were born in that seat — the seat where you get your hands dirty rather than manage — out of the simple wish to somehow cut down on the repetition. Today I lead a mid-sized (10–50 person) team as the design director of an MMORPG in active development, but the tools in this book did not come from a director's management kit; they came from a practitioner's hands. That is why the system in this book is not theory but a working environment that runs every day. The book also covers, in exactly the same way, one small puzzle game I built alone at home — quoting that game's actual git commits and code.

AI will not replace a game designer's job. What it does is free your hands from the busywork. What you do with those hands is still up to you. I hope this book serves as a practical guide to that transition.


How to Read This Book

You do not have to read this book in order from cover to cover. Pick the route that fits your situation. If terminals and installation are new to you, open to 1.0, "Before You Begin," before anything else — it is the chapter that eases the fear of the black screen first.

Route Path Who It Is For
The adoption route 1.0 (setup) → Part 1 (adoption) → Part 2 (information architecture) → one part for your own discipline Game designers just starting with AI tools
The full route Parts 1–2 → discipline parts (3–15) → process (16–19) → operations (20–24) Leads designing team-wide adoption
The indie/solo route 1.0 (setup) → Parts 1–2 → Part 23 (personal game development) → each chapter's "Solo Scale-Down" Developers building alone or as a hobby, with no team
The general-work route Parts 1–2 → Part 17 (meeting notes) → Part 16 (collaboration) → Part 18 (decision-making) → Parts 21–22 (self-improvement and governance) Designers, PMs, and office workers outside games
The problem-solving route Appendix index → backward to the relevant chapter Readers with a problem to solve right now

Every chapter ends with a "Try It Yourself" section. The goal is not a chapter you read and close, but one that gets your hands moving at least one step in your own environment, today.

One more word for readers who work outside games. Many of this book's workflows — turning meeting notes into decisions, tracing a decision's ripple effects, verification gates (this book's term for a quality gate where a human or a checker verifies output), cost management, copyright and ethics — work as is, with no connection to games. Feel free to read "game design" as your own job. The "Beyond Games" box in each chapter is the bridge, and if you are short on time, the 90-minute express course (17.1 → 16.2 → 22.1 → 21.1) alone will let you feel the core skeleton with your own hands.

Let me draw one distinction here. In this book, "solo" is used in two senses. One is the solo director — a lead who carries several people's worth of collaboration context alone. The other is the solo or hobbyist developer making a game by themselves. The "Solo Scale-Down" at the end of each chapter is for the latter: it lays out how to take just that chapter's core with no team and no company folder.

You can read the book straight through as a single volume, or split it into two halves — Parts 1–15, "foundations and disciplines," and Parts 16–24, "process and operations" — and start with whichever you need. And most of the code in this book runs as is with no external dependencies, using only the Python standard library. Only a few tools, such as the relation graph, need packages beyond the standard library (networkx, PyYAML), and in those spots a one-line install (pip install …) sits right next to the code. Apart from those cases, there is nothing extra to download — you can copy a code block and run it right away to see for yourself.

If a term stumps you, do not stop there — move on for now. If the black terminal feels foreign, that is not a flaw in the tool but a matter of familiarity, and Chapter 1.0 and Part 1 will close that distance with you.


The Easiest Way to Use This Book

Finally, a tip on the fastest way to put this book to work: feed the book itself, whole, to an AI tool like Claude Code.

This book was not written for human readers alone. The full prompts, code, and verification procedures in each chapter are written in a form an AI can understand and reproduce directly. So, in your own project folder, you can hand this book to an AI — as a PDF or as plain text — and ask something like, "Read the consistency-check patterns in this book and build a checker that fits our data sheets." The AI will then set up that chapter's workflow adapted to your environment. Both routes are open: building along chapter by chapter with your own hands, or handing the whole book to an AI and building together.

One thing does not change, though. The final call on what to adopt and what to reject is — as this book repeats from beginning to end — still yours. Even if you have an AI read the book and install the system, the seat where a human reviews what that system produces stays occupied by a human. Even the easiest way to use this book runs on that principle.


One Promise

No table in this book contains a number inflated to persuade you. Instead of overstating the impact, I show you the structure that produces it. Move the same structure into your own project, and you will measure your own numbers. That is the most honest help this book can offer.


The Promise Behind the Word "Reconstruction"

Throughout this book, the outputs in worked transcripts often carry the label "reconstruction" — for example, "Step 3 — Claude's Output (Reconstruction)." Let me make one precise promise, just once, about what that word preserves and what it touches. A book that puts honesty on its cover must not leave this spot, of all spots, blurry.

"Reconstruction" does not mean invented; it means an actual session, edited. The boundary runs like this.

Preserved Verbatim Edited
The full prompt I typed — in a form you can copy and use directly Company, project, NPC, and teammate proper names → book-safe anonymization (IP protection)
The structure and the failures of the AI's output — the off-target candidates, the spots where it quietly bent a rule, the back-and-forth where I rejected and reran it Length — side branches that do not fit the text are trimmed into "excerpts"
Code, constants, and verification values — kept as is, so they can be reproduced by running them yourself Line breaks, spacing, and other typesetting adjustments for the page

Put another way, no reconstructed output adds an inflated number or a success that never happened. It was anonymized, excerpted, and tidied for the page — nothing more. Nowhere has a failed output been rewritten as a success; if anything, I deliberately left the failures in, because they are what show you what a human rejects. (By contrast, anywhere code execution results or system logs are marked "measured" or "quoted verbatim," they were carried over without editing.)

Translator's note: In this edition, the prompts and outputs inside code blocks have been translated into English for readability. All worked transcripts in this edition are translations of the original Korean sessions — not re-runs performed in English. The verbatim Korean originals — the artifacts this book promises not to retouch — are preserved as is in the Korean edition. Code syntax, identifiers, numbers, and verification values are untouched. Currency conversions are approximate, at roughly 1,500 won to the US dollar as of mid-2026. Where the Korean text itself is the subject — dialogue registers, honorifics — or where a transcript is explicitly quoted verbatim, the original Korean is kept with an English gloss.



Contents

Part 1 · Foundation

1.0 Before You Start — Install, Account, Pricing, and the Terminal Survival Kit

Section 1.1 is the "first encounter." It is where you sit down in front of a blinking cursor and try typing something. But before you can take that seat, a few things need to be in place. The tool has to be installed, you have to be logged in, you should roughly know how the billing works, and you need to be able to type a few characters into a black screen. This chapter sits one step before 1.1.

Most introductory books skip this stage. They write a single line — "open your terminal" — and move on. But that single line is exactly where beginners get stuck. Where is the terminal? What do I need to install? What do I do when red text appears mid-install? Someone who stalls on the first line never reaches 1.1. This chapter has exactly one goal: making sure you do not get stuck on the first line.

The chapter has five parts: installation, account and login, pricing concepts, a terminal survival kit, and a "5-minute first run" checklist. Follow them in order and you will be ready to take the seat in 1.1.

flowchart LR
    A["1.0 Preparation
(this chapter)"] --> B["1.1 First encounter
(sitting at the cursor)"] A1["① Install"] --> A2["② Account & login"] A2 --> A3["③ Pricing concepts"] A3 --> A4["④ Terminal survival kit"] A4 --> A5["⑤ 5-minute first run"] A5 --> B classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; class A1,A2,A3,A4,A5 human class B pass

1.0.1 Install — One Line per OS

For installation, the rule is to follow the official instructions. Tools change often, and installer files from unofficial sources are dangerous. So this book does not print download links; instead, it shows you how to find the official path. Type "Claude Code official docs" or "Claude Code install" into a search engine and Anthropic's official documentation page comes up first. The safest move is to use the install command from that page exactly as written.

Still, it helps to know the big picture. Claude Code — this book uses the official English name throughout — is a tool that runs in the terminal, and it usually installs with a single one-line command. The flow differs slightly by OS.

OS What you need Install flow (conceptually)
Windows PowerShell (built in) Paste the one-line install command from the official docs into PowerShell
macOS Terminal (built in) Paste the one-line install command from the official docs into the terminal
Linux Terminal Paste the one-line install command from the official docs into the terminal

The flow is the same on all three: open the terminal → paste the one line from the official docs → press Enter. There is no need to memorize commands. Copying and pasting from the official docs is the standard way.

If red text (an error) appears during installation, do not panic. The install errors beginners run into are almost always one of two things: a permissions problem, or a missing prerequisite tool (a runtime like Node.js, for example). If red text appears, copy the whole message verbatim and search for it, or ask an AI — nine times out of ten that solves it. An error message is not an enemy; it is a clue.

How to check whether the install worked: type claude --version in the terminal and press Enter. If a version number appears on one line, the install succeeded. If you get something like "command not found," either it is not installed yet or you need to open a fresh terminal. Close the terminal completely, open it again, and check once more.


1.0.2 Account and Login

Finishing the install does not mean you can use it right away. Claude Code is a tool that borrows Anthropic's AI models, so there is a login step to confirm who is using it.

The flow is simple. Run claude in the terminal for the first time and a login prompt appears. Usually a web browser opens automatically, and you log in there with an Anthropic account (if you do not have one, you can create one on that screen). When the login finishes, the browser shows something like "you can return to the terminal now," and the terminal side shows a completion message as well.

Two spots trip up beginners here.

First, the browser may not open automatically. In that case the terminal prints a long address (URL) on one line. Copy that address, paste it into your browser's address bar, and go. You are not blocked — you just do one extra step manually.

Second, account types can be confusing. How the account you used in the web chat (Claude.ai) connects to Claude Code's account and billing may differ depending on the policy at the time. The most accurate guides are the login screen itself and the official docs. Follow what the first-run screen tells you and the login usually goes through without trouble.

Once you have logged in, it stays logged in on that PC. There is no need to do it every time.


1.0.3 Pricing Concepts — Flat-Rate Subscription vs. Pay-As-You-Go API

The part beginners worry about most is "how much is this going to cost?" There is a vague fear that every character you type adds to the bill. Getting the big picture first shrinks that fear. Billing comes in two broad flavors.

Method How you are charged Analogy Who it is for
Flat-rate subscription Fixed monthly amount A flat-rate phone plan Beginners, everyday use
Pay-as-you-go API By usage (per token) An electricity meter Bulk work, automation, integrations

A flat-rate subscription means paying a fixed monthly amount and using the tool up to a limit. It is like a flat-rate mobile plan. The same amount goes out every month, so it is predictable, and you do not have to think about "how much does each line cost." That is why beginners usually find it easier on the nerves to start with a flat-rate subscription (author's estimate — exact plan tiers and limits change over time, so check the official pricing page). If you hit the limit, you wait for the next cycle or move up to a higher plan.

Pay-as-you-go API billing charges in proportion to actual usage (tokens). Like an electricity meter, you are billed for what you use. It suits bulk processing, automation pipelines, and integration with other programs. Used with care it is efficient, but at the beginner stage, before you have a feel for your usage, costs can be hard to predict.

What a token is and why billing is based on it is covered in detail in 1.2 (AI models, tokens, and the harness). For now, remember just one thing: beginners usually start with a flat-rate subscription. The monthly amount is fixed, so you can practice without the fear of "what if I get hit with a surprise bill." Plan names, prices, and limits change often, so this book does not print specific numbers. The contents of this book were written as of mid-2026, and pricing, models, and features keep changing after that. The official pricing page is the most accurate source for current values.

One-line summary: the fear that money drains every time you type → with a flat-rate subscription, it is fixed every month. Starting with flat-rate keeps a beginner's mind at ease.


1.0.4 Terminal Survival Kit — Making the Black Screen Less Scary

Now for the biggest wall: the black screen. This is why 1.1 opens with "you freeze in front of the blinking cursor." To hands that have worked in GUIs for 24 years, the terminal feels foreign. But the number of commands you need so you do not stall on the first line is small. The six below are enough.

Command Read as What it does Analogy
pwd "p-w-d" Shows which folder you are in right now "Where am I?"
ls "l-s" Lists what is inside the current folder Opening a folder window
cd 폴더이름 "c-d" Moves into that folder (폴더이름 is Korean for "folder name" — replace it with the actual folder name) Double-clicking a folder
cd .. "c-d dot-dot" Moves up one folder level The Back button
Enter "enter" Runs the command you typed The OK button
Ctrl + C "control-c" Stops whatever is running right now The Stop button

(Windows PowerShell accepts ls, cd, and pwd as-is. So do macOS and Linux. That is why these six work regardless of OS.)

Here is what these six do, as a picture. Moving around in the terminal is ultimately just stepping in and out of folders — the same motion as double-clicking a folder or going back in a GUI.

flowchart TD
    Q["pwd
Where am I?"] --> L["ls
What is in here?"] L --> D{"See a folder
to enter?"} D -- "Yes" --> IN["cd folder-name
Go in"] D -- "No, go up" --> UP["cd ..
Step out"] IN --> L UP --> L RUN["After typing a command"] --> ENT["Enter
Run"] STUCK["When something seems stuck"] --> STOP["Ctrl + C
Stop and get the cursor back"] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef human fill:#fde68a,stroke:#b45309,color:#000; class Q,L,IN,UP,ENT,STOP code class D human

The real reason the black screen feels scary is the sense that "one wrong keystroke will break something." But none of the six commands above breaks anything. pwd, ls, and cd only look or move; they never delete or change files. Enter only runs, and Ctrl + C only stops. So feel free to type these six anytime, with peace of mind.

Sometimes the screen looks frozen. You typed a command and nothing happens for a while, or the cursor blinks on a different line as if waiting for something more. Press Ctrl + C once and you usually get the original prompt back. Just knowing this "stop button" exists makes the black screen far less scary. If you get stuck, escape with Ctrl + C and start again.

Finally, when typed characters pile up and the screen gets noisy, you can clear it. Windows PowerShell, macOS, and Linux all clear the screen with the clear command. Clearing does not undo anything you did; it only tidies up what is visible.


1.0.5 The "5-Minute First Run" Checklist

If you have come this far, the preparation is done. Pass the five boxes below within 5 minutes and you have earned the seat in 1.1. If any box blocks you, go back to the matching section (1.0.1–1.0.4).

flowchart LR
    C1["① The terminal
opens"] --> C2["② claude --version
shows a version"] C2 --> C3["③ Run claude →
login complete"] C3 --> C4["④ pwd & ls
show my folders"] C4 --> C5["⑤ Ctrl+C
gets me out"] C5 --> OK["✅ On to 1.1"] classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; class C1,C2,C3,C4,C5 human class OK pass

With all five boxes checked, the black screen is no longer an unknown wall. The tool is installed, you are logged in, you know the big picture of how billing works, and you can move around the screen and stop things. 1.1 starts on top of this preparation. Go take the seat in front of the blinking cursor and, for the first time, type "summarize what's in this folder."


1.0.6 Python and pip — To Run the Tools (Only When Needed)

The early parts of this book (Parts 1 and 2) can be followed with natural-language prompts alone. From Part 4 on, however, some chapters run small Python scripts directly (e.g., pip install pyyaml, pip install pyvis). If Python is new to you, that is fine. There are two paths.

First, install it yourself. Download Python from python.org and install it (be sure to check "Add to PATH" on the install screen), then confirm with python --version in the terminal. pip is the package installer that ships with Python; you fetch the packages you need with one line, like pip install pyyaml.

Second, let the AI handle it (recommended). The easier path is to delegate the environment setup itself to the AI. Ask in the terminal like this.

Check whether Python is installed, and if not, tell me how to install it for my OS.
Then give me a one-line command to install the pyyaml package this chapter needs.

(The prompt above, kept in the original Korean, says: "Check whether Python is installed, and if not, tell me how to install it for my OS. Then give me a one-line command to install the pyyaml package this chapter needs.")

The AI inspects your environment and produces the install commands for you. If you get stuck, paste the error message right there and ask "how do I fix this error?" One round of this pattern per tool-running chapter is enough. At any stage where Python and pip feel like too much, that chapter's "Solo Scale-Down" shows a lighter path that goes without code.


Next Chapter Preview


Try It Yourself

setup 1. Open the terminal for your OS (Windows: PowerShell, macOS: Terminal). 2. Search for "Claude Code official docs" and keep the official install page open. 3. Set a timer for 5 minutes — the goal is to pass the five boxes of the 1.0.5 checklist.

prompt (type one line at a time, in order — these are commands, not natural-language questions; the Korean comments say: ① version shown = install succeeded, ② which folder am I in, ③ what is in this folder, ④ go up one level (then ls again), ⑤ run Claude Code (follow the login prompts if they appear))

① claude --version      # version shown = install succeeded
② pwd                   # which folder am I in
③ ls                    # what is in this folder
④ cd ..                 # go up one level (then ls again)
⑤ claude                # run Claude Code (follow the login prompts if they appear)

verify - If ① shows a version number on one line, the install is done. If you get "command not found," close the terminal, open it again, and try once more. - While looking around and moving with ②, ③, and ④, confirm for yourself that nothing breaks. These three are safe commands that only look and move. - If ⑤ seems to hang, get out with Ctrl + C. If you got out, you have confirmed with your own hands that "there is a stop button."

Solo Scale-Down

If you are an individual with no team and no company folders, start by learning just the install (①) and escaping with Ctrl + C. Confirm "the tool is installed" with claude --version and "I can get out even when stuck" with Ctrl + C, and half the fear of the black screen is settled on your own within 5 minutes. For pricing, start with a flat-rate subscription and you can practice as much as you like without worrying about cost.

1.1 A Game Designer's First Encounter with Claude Code

A black screen comes up. The cursor blinks. In front of it sits a game designer with 24 years in the industry. Hands that have spent 24 years working in PowerPoint and Excel, wikis and Figma, pause for a moment over the keyboard. It looks a little like DOS, from the days of after-school computer classes. Since then, a terminal was something you only ever saw on a programmer's desk. I don't know what to type, and it feels like typing the wrong thing will break something. That hesitation is where this book begins.

Most people close the window right there. Then, in the next meeting, they go back to repeating "we really should be doing something with this." This chapter sits with you through the first 30 minutes without closing that window. The goal is not a grand adoption strategy — it is to put into your hands what to type at that blinking cursor so the distance starts to dissolve.


When a game designer first sits down in front of an AI coding tool like Claude Code, fascination and discomfort arise at the same time, in the same seat. The fact that these two feelings collide is itself the first clue to adoption.

The fascination has clear reasons. A data-sheet consistency check that used to eat half a day finishes in minutes; sprawling meeting minutes get summarized into a table of decisions; a game design document (GDD) buried a year ago can be pulled back up with one line of natural language.

The discomfort has equally clear reasons. The black screen, the blinking cursor, the English commands — none of it looks like the everyday workspace. A game designer's day flows on top of GUIs, and typing characters into a black terminal does not sit well with the job's identity. But this discomfort is not a defect in the tool; it is the adaptation cost of a person used to GUIs. Just acknowledging that point clears half the distance already.

This book is about closing that distance. Section 1.1 sits with you at the first encounter and sorts out what to look at, what to try, and what can safely wait.


1.1.1 Why Now, and Why Game Designers Should Use AI

Game design joined the AI wave later than other roles. People who handle code went first; artists and visual designers came next. Game designers commonly fell into the pattern of repeating "we should try something" while putting it off.

The reasons for putting it off were rational. A game designer's output is not structured the way code is; it is a mix of text, tables, diagrams, meetings, and verbal agreements. AI output looked unreliable, plausible lies were dangerous, and it was doubtful whether AI really understood game systems.

But between 2024 and 2026, three things changed.

First, the reasoning ability of AI models crossed a threshold. They go beyond simple sentence generation to handle complex system design, consistency verification, and impact analysis. Recent Claude-family models can assist with a substantial share of a game design workflow. That does not mean you can hand all of it over: verification and responsibility still belong to humans. (How wide the assistance goes varies greatly by task type and team maturity — author's estimate, unverified.)

Second, the harness matured. A tool like Claude Code is not simple chat. It reads and writes files directly, runs commands, and takes the results back as input. It resembles the way a person works.

Third, operational techniques like memory, atoms, and skills have settled in. Instead of using AI once and throwing the session away, there is now a methodology for accumulating a team's knowledge so the system gets smarter over time. That accumulation is exactly what the later parts of this book cover.

Put these three together, and this becomes a rational moment for game designers, too, to adopt AI. It pays to start before falling further behind.


1.1.2 What Claude Code Is — How It Differs from Other Tools

The AI tools a game designer usually encounters come in two kinds: chatbot-style tools where you type questions into a chat window (ChatGPT, the Claude web app), and editor-integrated tools that autocomplete inside a code editor (Cursor, Copilot).

Claude Code is a third kind. It runs in the CLI (terminal) and has access to the entire environment a person works in. Here, on one page, is where the three kinds diverge.

Chatbot Editor-integrated Claude Code Input location Web chat window Code editor Terminal File access Upload required Open files Entire project Command execution Not possible Partial Free (within permissions) Output form Text Code suggestions File changes & run results Design-work fit Low Low High Analogy Front desk Autocomplete pen Next-desk colleague

A game designer's work is not code; it is documents, tables, and relationships. Claude Code's strength is that it sees, understands, and manipulates the entire folder a person works in. You do not have to copy and paste material into a chat window every time, the way you do with a chatbot.

In office terms, the chatbot is the front desk. One question, one answer, and you have to bring your materials out again every time. Claude Code is closer to the colleague at the next desk. It knows where the materials are, opens the files with its own hands, and puts the organized result back on your desk. The same Claude does different things well depending on which desk you seat it at.


1.1.3 The First 30 Minutes — What to Look at, What to Try

Installation and setup are covered in 1.0. Section 1.1 focuses on what to experience in the first 30 minutes to shrink the distance. The first 30 minutes split into four segments, and each segment dissolves another piece of the distance, step by step.

flowchart LR
    A["Look (5 min)
Summarize this folder
→ 'It understands natural language'"] B["Search (5 min)
Find keywords, recent files
→ 'It handles my daily chores'"] C["Convert (10 min)
Minutes→table, Excel→Markdown
→ 'It creates time'"] D["Verify (10 min)
Doubt the output just made
→ 'This tool must be doubted'"] A --> B --> C --> D D --> E["80% of the distance dissolved"] classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; class A,B,C ai class D human class E pass

The First 5 Minutes — Look

Open a terminal and run Claude Code: a black screen appears and the cursor blinks. Many people stop right here, but stopping means missing the first chance to dissolve the distance. Type one sentence in natural language.

Summarize what's in this folder

One sentence — "Summarize what's in this folder" — neither code nor a command. Claude Code scans the folder structure and answers with a summary of what files are there and what kind of project this appears to be. Type this one line in a folder of design documents and a response comes back roughly like this.



Summarize what's in this folder

I scanned this folder. It looks like a working folder for game data and design documents. File type Count Estimated contents ───────────── ──── ─────────────────────────── .xlsm 12 Data sheets (items, quests, NPCs, etc.) .md 34 Design docs and meeting-minute drafts .png 88 UI mockups and reference images Subfolders 6 /data /docs /ref /meeting /export /tmp Observations: - The /meeting folder has 5 meeting minutes from the past week piled up. - item_table.xlsm and quest_table.xlsm appear to share the same ID column (an FK candidate). - The /tmp folder holds many temporary export files — possible cleanup targets. What would you like to look at first?

(The response, in short: it scanned the folder, identified it as a game-data and design-doc workspace, tabulated the file types — 12 .xlsm data sheets, 34 .md docs and minutes, 88 .png mockups, 6 subfolders — then observed that five recent meeting minutes have piled up, that item_table.xlsm and quest_table.xlsm appear to share an ID column (an FK candidate), and that /tmp holds stray export files; it ends by asking what to look at first.)

From this single response, the fact that this tool understands natural language, not commands, sinks in physically. Receiving an organized picture of what is where, without opening the folder yourself, cuts half the distance on the spot.

The Next 5 Minutes — Search

Give it a small search. Requests like "show me the 5 Markdown files modified in the last week," or "where are the files that mention the word 'reward'?" The work a game designer used to do daily — opening Find in Files, typing a keyword, scrolling through results — Claude Code takes in natural language and returns organized. The sense that this tool handles your daily chores settles in during the second five minutes.

The Next 10 Minutes — Convert

Give it a small conversion. Requests like "extract only the decisions from these meeting minutes into a table," or "turn rows 1–10 of this Excel sheet into a Markdown table." A conversion that takes 5–10 minutes by hand getting done in moments turns the earlier sense — the tool replaces search — into a new one: the tool creates time.

The Last 10 Minutes — Verify

Stopping here is dangerous, because AI tells plausible lies. Make it doubt the output it just produced.

Check again whether the table you just made matches the original. Anything missing?

The prompt says: "Check again whether the table you just made matches the original. Anything missing?" Claude Code then rereads the original, compares it line by line against its own output, and reports the mismatches. In practice, this segment often produces a screen like this.



Check again whether the table you just made matches the original. Anything missing?

I re-checked the original (meeting_0603.md) against the table I just made. I found 3 mismatches. ⚠ Omitted "Tutorial reward cut by 50%" — in line 12 of the original but missing from the table ⚠ Mistyped Owner "Kim OO" → entered in the table incorrectly as "Lee OO" ⚠ Ambiguous "Rediscuss rewards" — unclear in the source whether it is a decision or a hold. I placed it under 'decided' on my own judgment; it needs confirmation. Shall I rebuild the corrected table? Tell me how to handle the ambiguous item and I will reflect it.

(In short, it found 3 mismatches: an omission — "tutorial reward cut by 50%" is in line 12 of the original but missing from the table; a wrong entry — owner "Kim OO" mistyped as "Lee OO"; and an ambiguity — "rediscuss rewards" is unclear in the source as decision or hold; the AI says it placed it under 'decided' on its own judgment and asks for confirmation, offering to rebuild the table.)

The point of the last 10 minutes is that the tool knows how to doubt its own output — and that it does the doubting together with a person. The attitude in that third item, asking back "I made a judgment call here, please confirm," is the safeguard that keeps verification in human hands.

After 30 minutes, 80% of the distance is gone. The remaining 20% shrinks slowly across the chapters that follow.


1.1.4 Scenes from the Office — Six Months on a Mid-Size Team

On an MMORPG project I run as design director (hereafter "Project A"), the design team (4–5 people) has been running a Claude Code–centered workflow for about six months (Project A's full development team is mid-size, 10–50 people). Here are a few of the scenes.

Let me pull out the single most tangible case as an actual measurement: FK (foreign key) consistency checks across 30-odd data sheets. The task is tracing, by eye, whether the IDs in one sheet are referenced correctly in the others — and as sheets multiply, the combinations grow at a compounding rate.

Half a day to 5 minutes. I will not generalize from this one line. Other tasks save less, or pick up new review time instead. The other scenes from the same six months I record only as directions and ratios.

My felt sense is that once each of these tools was built, the time it saved over six months accumulates in person-months, not person-weeks (the exact total is unmeasured — an estimate). That time went into deeper design work.

A tool like this works for a long time once built. But 'long' does not mean 'unattended.' It lasts only when an operator and a verification structure come with it; leave the tool in place and let the person walk away, and it rots within two quarters. The later parts of this book cover how each of the tools above gets built and operated.


1.1.5 Fears and Expectations — Dealt with Honestly

Let's be honest about the fears game designers commonly bring to an AI tool. Facing them instead of looking away is the first step of adoption.

The fear that "AI will replace my job" is half right and half wrong. AI does replace the simple chores — consistency checks, document conversion, search — but it cannot replace decisions, priorities, or the design of player emotion. If anything, the game designer who uses AI well is freed from chores to focus on the essentials. Ask yourself: "In my job, what is the ratio of chores to essence?" If chores are 70%, the 30% of essence stays yours — and the point is that that 30% becomes more important.

The question "who takes responsibility when AI is wrong" comes up just as often. Responsibility for a game designer's decisions always belongs to the game designer. Using AI output as-is without verification is the designer's mistake, not the AI's. Designing the verification procedure alongside the tool is part of adoption. The 'make it doubt its own output' move from the last 10 minutes of 1.1.3 is the smallest seed of that procedure.

The fear "I can't use it because I don't know code well" resolves quickly. Claude Code runs on natural language, so you can start without knowing code, exactly as you are. Because you read the scripts the AI writes together with it, after a few months you find yourself reading and modifying simple scripts. The learning comes along on its own.

"The tools change too fast" is another common worry. Trying to keep up with every model, feature, and trend wears you out. Learn only the one or two features that help your own workflow in depth, and look at the rest when you need it.

One thing to say in advance: even if the response screens in 1.1.3 looked smooth, the real first 30 minutes will mix in answers that miss, summaries of the wrong files, and output that stalls. That is normal. This book is not a smooth success story; it spends more pages on how to re-ask and correct output that went wrong.


1.1.6 How to Use This Book

This book is organized into 24 parts. You do not have to read it front to back. Pick whichever of the three patterns below fits your situation.

Pattern Path Time
Adoption pattern Part 1 (adoption) → Part 2 (information architecture) → 1 part for your own field 1–2 months
Full pattern Parts 1–2 → field parts (3–15) → process (16–19) → operations (20–24) 6 months–1 year; suited to teams
Problem-solving pattern Appendix index → backward to the relevant chapter about 1 week, when you have a problem right now

If you have no idea which path to pick, split by where you stand. If you work outside games — in a planning or product role, as a PM, or as a general office professional — instead of the three patterns above I recommend the "General-Role Path" (Parts 1–2 → Part 17 meeting minutes → Part 16 collaboration → Part 18 decision-making → Parts 21–22 self-improvement and governance) — the core skeleton stands intact even when you skip the game-domain chapters, and each chapter's "Beyond Games" box is the bridge for carrying it over to your own role (index in Appendix F.5). If you are short on time, following just four chapters — 17.1 → 16.2 → 22.1 → 21.1 — is enough. If you are a non-technical reader just starting out with AI tools, take the 'adoption pattern,' hold on to one field (your own, or the closest one) all the way through, and for the deeper field parts (4, 8, 11, and the like) just pick up the 'one line for non-specialists' at the start of each, going down into the body only when needed.

Every chapter in this book stops short of academic depth, at the level you can actually operate. The goal is to carry over, as-is, techniques that really ran for six months on a mid-size team, and to walk together the path of starting small and growing big.


1.1.7 Into the Next Chapter

1.1 was a chapter about shrinking the distance. 1.2 steps one level inside and explains the tool's basic mechanisms in language friendly to game designers. The goal of 1.2 is to make words like model, token, context, and harness nothing to fear. The full setup (memory, permissions, settings.json) comes in 1.3.


Key Takeaways

Next Chapter Preview


Try It Yourself

setup 1. Open a terminal (PowerShell on Windows, Terminal on macOS). 2. Move into a folder where your design documents live, then run Claude Code (installation: 1.0). 3. Set a timer for 30 minutes — 5 min (look), 5 min (search), 10 min (convert), 10 min (verify).

prompt (one line per segment, in order)

① Summarize what's in this folder
② Where are the files that mention the word 'reward'?
③ Extract only the decisions from these minutes into a table
④ Check again whether the table you just made matches the original. Anything missing?

(In order: ① "Summarize what's in this folder" ② "Where are the files that mention the word 'reward'?" ③ "Extract only the decisions from these minutes into a table" ④ "Check again whether the table you just made matches the original. Anything missing?")

verify - For ①, check with your own eyes that the folder summary matches the actual folder. - If ④ reports even one mismatch, that is a success. It means you watched the AI doubt its own output, live. - An answer that misses is not a failure either. Re-asking — "that answer was wrong, look at just this file again" — is part of the first 30 minutes of practice.

Solo Scale-Down

If you have no team and no company folder, pick any working folder on your own PC (say, your Downloads folder, or a notes folder) and try just prompts ① and ④ above. Confirm "it understands natural language" with ① and "it can doubt its own output" with ④, and you will have felt this chapter's two core points by yourself, within 5 minutes.

1.2 Model, Token, Harness — The Path a Task's Tokens Travel

It starts the moment I finish a task and look at the usage. Five meeting notes from this week have piled up in a folder, and before Monday morning's stand-up I have to distill them into a single page of "only what was decided." I type one line into the Claude Code window: "Pull only the decisions from the meeting notes in this folder and make them into a table." About 0.4 seconds after I press Enter, small gray text flickers near the bottom of the screen.

Reading meeting-2026-05-25.md ... (1,840 tokens)
Reading meeting-2026-05-27.md ... (2,310 tokens)

That gray text is the subject of this chapter. I toss in a single sentence in Korean; the tool chops it into tokens, reads files into tokens and feeds them to the model, takes the model's answer, and writes it to a file. Every time this round trip runs, a cost is charged, and material piles up in the model's "field of view." This chapter breaks down what happens behind that gray text in a game designer's language. Four words are enough: model, token, context, harness.

Terminology Notes - Model: the brain that produces answers. It comes in kinds with different sizes and characters, such as Opus, Sonnet, and Haiku. - Token: a small chunk of chopped-up text. Billing, speed, and the model's field of view are all counted in this unit. - Context window: the maximum number of tokens the model can hold in its head at once. - Harness: the chassis that puts the model to work. Claude Code is one example.


1.2.1 The Harness Loop — What That Gray Text Really Is

The gray text above is not random logging; it is one box in a fixed cycle. What a harness does, in the end, is spin the same loop fast. It reads files and feeds them to the model; when the model says "run this command," it runs it; and it feeds the result back to the model. The loop keeps spinning until the task is done.

flowchart TD
    Start([Human: one-line instruction]) --> Read[Harness: read files and materials
convert to tokens] Read --> Inject[Harness: inject into context
+ add to cumulative tokens] Inject --> Model{Model: decide next action} Model -->|needs a command run| Exec[Harness: run shell commands and scripts] Exec --> Result[Feed execution results
back in as tokens] Result --> Inject Model -->|answer is ready| Write[Harness: write files and output] Write --> Verify{Verification passed?} Verify -->|fail| Inject Verify -->|pass| Done([Save result]) classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; class Start,Verify human class Read,Inject,Exec,Write code class Model ai class Result data class Done pass

In this diagram, the only boxes a human touches are the top one (instruction) and the bottom one (checking the verification result); the loop in the middle runs autonomously under the harness. The gray text flickered five times while the five meeting notes were being read because the Read → Inject boxes ran five laps. In a web chat, I would have had to open five files myself and copy-paste them. The harness takes over that labor — and that is the decisive difference that makes a chatbot and a CLI-style harness different tools.

Each lap of the loop adds to the cumulative token count in the Inject box. So without understanding tokens first, neither the cost nor the limits of this loop come into view. Tokens first.


1.2.2 Tokens — The Real Currency a Task Spends

A token is not a character; it is a chunk the model cuts text into. As a rule of thumb, English runs at about 4 characters per token and Korean at about 2 characters per token (this is an operational approximation, not an official conversion — the actual value is set by the model's tokenizer and varies by sentence). Twenty Korean characters including spaces come to roughly 10 tokens.

Let me follow that meeting-notes task in tokens (the figures below are a single measurement of one run of this task; they vary with the length of the notes and the summary, so read them for orders of magnitude and ratios, not as absolute values).

Step What Tokens (input) Tokens (output)
Instruction The one line "pull only the decisions into a table" \~25
Reading notes ×5 Body of 5 md files \~10,400
Classification rule injection 1 meeting-category atom (JIT) \~480
Model reasoning, table writing 12 decisions into a table \~1,600
Verification re-entry Re-query for 1 omission the linter caught \~320 \~210
Total \~11,225 \~1,810

Two things stand out. First, the instruction I typed is 25 tokens, but the task as a whole tops 11,000 tokens of input alone. Almost all of the cost comes not from my sentence but from the material the tool read in. Second, output (1,810) is about one-sixth of input (11,225). Most design automation reads a lot and writes a little, just like this. So if you want to cut costs, taming the volume of input material works far better than polishing the output.

After the task ends, typing /context shows how much of the context that session occupied. Unwatched, tokens drain away as unconsciously as printer paper; once made visible, your posture changes. That visibility is where saving starts.

What tames tokens is not an abstract spirit of thrift but concrete techniques for handling input material in small pieces.

  1. JIT injection — JIT (just-in-time): instead of loading all material up front, pull it in by keyword matching only when needed. The "\~480 tokens of classification rule injection" in the table above is an example. What came in was not the entire meeting-classification rulebook (thousands of tokens) but the single matched atom.
  2. Summary cache — for long documents, keep an AI-facing summary separate from the human-facing original. The AI reads the summary.
  3. Atom splitting — keep one decision per file (details in 2.2) and you can pull in exactly the pieces you need, saving tokens.
  4. Context cleanup — compact when a session runs long. Claude Code supports automatic compaction.
  5. Model selection — running a simple conversion on a big model makes the same tokens cost more. That is the next section's topic.

Of these, #1, JIT injection, is a device that actually runs in this book's working environment. When a line of input comes in, the inject_memory.py hook matches memory atoms by score, picks only the top few to inject, and never blocks the workflow even when it fails (implementation details in 1.3). The token-saving principle — "only the needed material, only the top few, fail quietly" — sits in that one file of code, as is.


1.2.3 Models — Same Chassis, Different Engines

A model is like a car engine: into the same chassis called Claude Code, you can fit different engines — Opus, Sonnet, Haiku. Swap the engine and the character of the work changes.

Model Matching — Depth vs Speed & Cost → Speed · low cost Reasoning depth ↑ Opus Large engine Design review · GDD synthesis Sonnet Mid-size engine Meeting notes · daily 80% Haiku Compact Simple sheet conversion

Mapped onto design work, it splits like this. Work that needs deep reasoning and consistency — reviewing a system design, or synthesizing a first draft of a game design document (GDD, the detailed spec) that pulls multiple sources together — goes to Opus. The everyday majority, like extracting decisions from meeting notes or writing daily summaries, goes to Sonnet. Work involving almost no judgment, like simple format conversion of data sheets, goes to Haiku. Running the earlier meeting-notes task on Sonnet followed this rule — picking out decisions and moving them into a table calls for balance and speed more than deep reasoning.

There is a trap everyone falls into early on: the urge to run every task on the best engine, Opus. Follow that urge, and the cost and speed load comes back as an operational burden, and the sense of matching the model to the task never takes root. The real skill of operation is not picking a model in your head every time, but hardening the pattern into automation once it has settled.

These fixed choices are spelled out in settings.json or inside slash commands (details in 1.3). Harden them once, and the chore of choosing every time disappears.

Models ship a new version roughly every half year, and even under the same name, 4.5 and 4.6 are different. When a new version lands, I compare only the five core tasks of my workflow on identical inputs. Try to test everything and you wear out. The differences across five results are enough to decide whether to switch.


1.2.4 The Context Window — The Ceiling the Loop Fills Toward

I said tokens accumulate in the Inject box with every lap of the loop. The ceiling that accumulation runs into is the context window: the maximum number of tokens the model can handle at once. In human terms, it is working memory.

The earlier meeting-notes task accumulated input in the 11,000-token range — about 6% of the 200K ceiling, with plenty of headroom. But if you never switch tasks and drag one session on in the same window, you approach the ceiling. When it fills, old content gets cut, the model starts losing its "memory" of the earlier parts, and automatic compaction kicks in, replacing the previous conversation with a summary.

Four habits keep this ceiling under control.

Pattern When
Session split Open a new session when moving to a different topic
Explicit compaction When one task ends, compact down to the essentials
Memory externalization Move frequently used material out into atoms and JIT-inject it as needed
Context visibility Check current usage with /context

The heavy case game designers run into most often is work that needs meeting materials, design docs, and data sheets all at once — that is where the 1M option is useful. But 1M carries cost and speed burdens, so 200K is enough day to day; bring out 1M only when the bundle of material is genuinely large.


1.2.5 When the Harness Actually Filters Out Falsehoods — The Verification Box

Back to the Verification passed? box at the bottom of the loop diagram. Without this box, the model's plausible lies get saved straight to files. A model sometimes presents a confidently wrong answer as if it were correct (a hallucination), and while the frequency drops with each generation, it never reaches zero. So verification stays in place as a permanent box.

In design work, the dangerous hallucinations are specific: citing a data sheet column that does not exist, computing balance with the wrong formula, or summarizing something never decided in the meeting as if it had been decided. In the meeting-notes task, the third is the scariest — an item that was "only discussed and put on hold" quietly climbing into the decisions table.

There are five verification patterns.

  1. Source cross-check — compare the AI output against the original material again ("check whether this was really in that document").
  2. Round-trip conversion — convert A→B, then convert B→A back, and check that they match.
  3. Sample review — a human directly checks 3–5 random items from the output.
  4. Linter automation — automatically check whether the output violated the set format, ranges, or rules.
  5. Two-model cross-check — Opus reviews Sonnet's output.

You do not run all five every time; pick one to three based on the risk level of the task. For the meeting-notes task, I paired #4, a linter ("does every decision have an owner, a description, and a deadline?"), with one round of #3, sample review. The last row of the earlier token table — "\~320 tokens of verification re-entry" — is exactly that round trip where the linter caught an omission and asked the model again; in the loop diagram, one extra lap through fail → Inject.

If a human verified everything every time, the payoff of adoption would be cut in half, so verification itself is an automation target. For meeting-decision extraction, the linter checks for format omissions; for data sheet conversion, row counts, sums, and foreign key consistency; for GDD auto-generation, missing core sections. Whatever passes, no human needs to look at; only what fails gets looked at. Picture a cabinet stuffed with paperwork where you pull out only the folders with red tags. Designing things so the human gaze lands only where the danger is — that is the purpose of verification automation.


1.2.6 Where the Four Words Tie into One Task

Now let me lay out that meeting-notes task from start to finish — how the four words string onto a single thread.

Box What happens Which concept
1 Run Claude Code in the meeting-notes folder Harness
2 The task is meeting-note analysis, so pick Sonnet Model
3 About 11K tokens total, within the 200K window — OK Token, context
4 The meeting-classification-rule atom is JIT-injected automatically Token (saving)
5 The model outputs 12 decisions as a table Model, harness loop
6 The linter flags 1 format omission → re-query and patch Verification (one extra loop)
7 Save and commit as weekly-decisions-2026-W21.md Harness

Done by hand, opening five meeting notes, reading them, picking out only the decisions, transcribing them, and fixing the format takes 30 minutes. Automated, it shrinks to 5 minutes, and within those 5 minutes the only work for human hands is skimming a verification sample once. Human time lands only where it is truly needed — checking that no deferred item wrongly climbed into the decisions. The point is not the 25 minutes saved but the change in where that gaze goes.


1.2.7 Common Misconceptions

"Opus is always better" is the most common. Ignore cost and speed and it is true, but Opus on simple work is waste. Per-task matching is the answer.

"You always need the 1M context" also comes up often. 200K is enough for most things, and 1M carries a burden, so save it for genuinely large bundles of material.

"Verification is something humans do" is half right. Most of it can be verified automatically, and humans focus on the rest.

"You don't need to mind tokens" holds up to a point for solo work, but once several people use the tool together, the cumulative cost grows fast. Settling visibility and saving patterns from the start is the safer path.

"Harnesses make no difference" is surprisingly common too. The same model splits into what feel like different tools depending on whether it sits in a chatbot or a CLI. The presence or absence of the copy-paste labor we saw earlier is that difference.


1.2.8 Try It Yourself

Run this chapter's four words yourself on one small task.

setup

prompt

From the notes in this folder, pick only what was "decided"
and build a 3-column table of owner, description, and deadline.
Exclude deferred or still-under-discussion items, and for each row
in the table, also note the file name it came from.

(The prompt says: from the notes in this folder, pick only what was decided and build a three-column table of owner, description, and deadline; exclude deferred or still-under-discussion items, and note for each row which file it came from.)

verify

Solo Scale-Down

If you are just starting with the tool, keep only two things from the above. First, hand over the notes folder whole — do not copy-paste the contents yourself (leave that to the harness loop). Second, always get the output table with file names attached, and open the original only for the suspicious lines. Model selection and token visibility can wait until you are comfortable. Not carrying material by hand, and checking output against the original — these two habits alone settle half of the adoption.


Key Takeaways

Next Chapter Preview

1.3 Memory, Permissions, and Settings Infrastructure

I opened a new session and typed, "Let's take a look at skill cooldown balance." Before I pressed Enter, a small line of gray text flashed by at the bottom of the screen: [memory injected: 2 atoms, 1,842 chars]. I had not opened a single file, yet the cooldown rules document I had pinned down the week before was already attached in front of the model's input. That is the first sign of a working environment with infrastructure in place. The moment I turn the tool on, the tool remembers me.

For that scene to happen, three things must already be in position: what the AI remembers (memory), what the AI can do without human approval (permissions), and the central switch that turns both on and off (settings.json). The initial setup takes an hour at most, and that hour comes back as time saved every day for the next six months. It is an investment that recovers almost all of itself.

This chapter is a walkthrough that unfolds, in order, the single line of settings.json I actually run on my personal PC, the inject_memory.py that line calls, and the _jit_manifest.json that script reads. Read to the end and you will be able to put your finger on exactly which file and which line make "memory gets injected automatically" happen.


1.3.1 settings.json — The One Line Where Everything Starts

Conclusion first. On my personal PC, what turns on automatic memory injection is one single block inside settings.json.

{
  "hooks": {
    "UserPromptSubmit": [
      {
        "hooks": [
          {
            "type": "command",
            "command": "python ~/.claude/hooks/inject_memory.py"
          }
        ]
      }
    ]
  }
}

Spelled out in plain language, the block says this: "Every time the event where the user submits a prompt (UserPromptSubmit) fires, run a Python script called inject_memory.py once." That is all. The AI is not being clever and remembering things on its own — a script that a human registered cuts in once every time input arrives.

settings.json is the central file that controls every behavior of Claude Code, and it splits into two layers.

The two are merged when applied. So I keep the hooks and permissions the team should share in settings.json, and the absolute paths and personal tool paths that exist only on this home PC in settings.local.json. The split avoids git conflicts when collaborating and prevents the accident of personal settings leaking into the team repository.

Besides hooks, a few other entries come up often.

There is one operating habit here that matters more than anything else. A single small typo in settings.json can keep the tool itself from starting. One missing JSON comma breaks the parse. So a backup before every change is mandatory. These backup files actually sit on my PC.

settings.json.bak_2026-05
settings.local.json.bak_2026-05

Take one copy with a date suffix and rollback takes one second. Managing it with git is even better. Like a spare set of old keys in a drawer — you never use it day to day, but a moment always comes, in front of a locked door, when you need it exactly once.


1.3.2 inject_memory.py — What Actually Happens Inside the Hook

Now we go inside the script that settings.json calls. This is the spine of the walkthrough. The code is barely 100 lines, but the heart of it is five actions.

flowchart TD
    A["Session: user submits a prompt\ne.g. 'Review skill cooldown balance'"] --> B["UserPromptSubmit hook fires\nsettings.json runs inject_memory.py"]
    B --> C["Load _jit_manifest.json\nmetadata for 17 atoms"]
    C --> D["Sort by score, descending\n→ try regex match per atom"]
    D --> E{"Any matched\natoms?"}
    E -->|none| Z["Inject nothing\nexit 0"]
    E -->|some| F["Select only top max_matches (3)"]
    F --> G["Truncate if combined total\nexceeds 6,000 chars"]
    G --> H["Attach atom bodies\nin front of user input"]
    H --> I["Model receives atoms + prompt\ntogether → responds"]
    D -.->|any exception| Z2["Swallow every exception\nexit 0 (never blocks the flow)"]
    classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545;
    classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764;
    classDef human fill:#fde68a,stroke:#b45309,color:#000;
    class A human
    class B,C,D,E,F,G,H code
    class I ai

Spelled out, the five actions go like this.

1) Read the manifest. The script first opens ~/.claude/projects/C--Users-user/memory/_jit_manifest.json. This file holds the metadata for each atom — name, path, matching regex, score. On my personal PC, 17 atoms are currently registered.

2) Sort by score, descending. Every atom carries a score value. The higher the score, the earlier the atom gets its matching attempt. When several atoms hit the same keyword, this score decides who gets priority.

3) Match with regex. The user's input string is checked against each atom's regex pattern. If the input contains "쿨다운" (cooldown), the atom carrying the pattern 쿨다운|cooldown|GCD fires. The comparison is case-insensitive.

4) Cut to at most three. No matter how many atoms match, anything beyond max_matches (3 in my environment) is dropped, keeping only the top three. On top of that, if the combined length of the selected atom bodies exceeds 6,000 characters, the script truncates. Two layers of ceilings serve as the safety device that keeps the input from bloating.

5) Exit 0 on every exception. This is the heart of the design. Whether the manifest is broken, the file has vanished, or a regex is malformed, the script quietly swallows the exception and finishes with exit code 0. If the hook dies with a nonzero code, the user's prompt itself can be blocked. The principle — "even if memory injection fails, never block the user's flow of work" — is written into the outermost try/except of the code.

The center of gravity sits on 4 and 5. Number 4 (the ceilings) keeps memory from exploding the token count, and number 5 (swallowing exceptions) keeps the infrastructure from getting in the way of the work. Both are two faces of the same philosophy: automation must never make itself a burden on the human.


1.3.3 _jit_manifest.json — The Keyword Dictionary That Wakes Atoms

In 1.2 I promised the token-saving principle — only the material you need, only the top few, fail quietly — and deferred the implementation details to this chapter. Those details live in the manifest that inject_memory.py reads. It is the heart of JIT (Just-In-Time — loading material only at the moment it is needed), and a single atom entry looks like this.

{
  "atoms": [
    {
      "name": "combat_cooldown_rule_v2",
      "path": "atoms/combat/combat_cooldown_rule_v2.md",
      "regex": "쿨다운|cooldown|GCD",
      "score": 80
    },
    {
      "name": "user_health",
      "path": "memory/user_health.md",
      "regex": "건강|복약|컨디션|약물",
      "score": 95
    }
  ],
  "config": {
    "max_matches": 3,
    "case_insensitive": true
  }
}

Four fields define one atom.

The max_matches: 3 in the config block is the source of the "at most three" ceiling we saw in 1.3.2. Edit the manifest by hand and the behavior changes immediately.

One note on scale. My personal PC runs light: 17 atoms, one manifest. The company production environment (Project A), by contrast, has 304 team atoms and 48 skills registered as of a May 2026 backup. One hot atom has a score that has climbed to 356.53 (the view_html_filename_convention family, which covers file-naming conventions) — it did not start high; the number is a trace accumulated through repeated calls and verification.

What the gap between 17 on a personal PC and 304 at the company tells you is this: the JIT mechanism is the same, but the speed and scale at which material accumulates is proportional to project density. There is no need to build 304 from day one. Start with five core atoms, pin down one or two more each week, and before long the manifest thickens on its own.

Author's estimate (unverified): the claim that score accumulates with match and verification counts is an interpretation based on operating patterns. The scoring formula itself varies with each environment's manifest design, so an absolute value like the 356.53 above is a measured snapshot of my environment, not a general standard.

This is also the place to restate the principle that memory is kept in two layers.

Layer Location When loaded Purpose
Global ~/.claude/memory/ Every session Your identity, collaboration rules, language settings
Project ~/.claude/projects/<project>/memory/ Sessions of that project Per-project atoms, rules, material

Keeping global light is the safer side. When global grows heavy, that weight accumulates as a token cost on every single session. In office terms, global is the business-card holder on your desk (the lighter it is, the handier for daily use), and project memory is the folders in the cabinet beside you (it can grow thick per project without weighing on your day). So the auto-loaded global holds only the essentials, while the rich material piles up in project memory and gets woken by JIT only when needed.


1.3.4 Permissions — A Whitelist That Accumulates the Traces of Your Work

Now the third axis of the infrastructure: permissions. Claude Code can delete files, run commands, and call external APIs. Power comes with risk. The permission system manages that risk.

Permissions split into two kinds: what runs automatically without human approval, and what requires approval every time. Which side holds what is defined in the permissions block of settings.json.

{
  "permissions": {
    "allow": [
      "Bash(ls:*)",
      "Bash(git status:*)",
      "Bash(git diff:*)",
      "Read(*)",
      "Grep(*)"
    ],
    "deny": [
      "Bash(rm -rf:*)",
      "Bash(git push --force:*)"
    ]
  }
}

A shift in perspective is needed here. This allow list is not a mere configuration value — it is a trace of accumulated work. At first it sits nearly empty, with only reads and searches auto-allowed. Then, after a month or two of repeating the same work, patterns emerge where you think "approving this command every single time is a chore," and you move them into allow one by one. The list that grows long is a fingerprint of what I have been repeating with this tool.

My company environment (Project A) holds about 80 auto-allow patterns. It started at 20, and 60 more attached themselves over six months — read those 60 backwards and they reveal what work got repeated over the past half year. Data sheet extraction, relation map generation, schema documentation — the tools you use often become the permissions you allow often.

Four patterns settle in for operating permissions.

When an approval popup appears every time, the human wears out. There are devices to cut the fatigue: bulk-registering frequent patterns with a slash command like fewer-permission-prompts, granting a temporary allow for a single session only, or — for personal work only — running a full auto-allow mode. I do not recommend that last option in a team environment, though.

The balance between fatigue and safety is yours to tune. Too strict and the work does not roll; too loose and accidents happen. Even if you start loose, with a quarterly cleanup cycle in place the balance settles on its own.


1.3.5 When a Session Starts — How Memory and Permissions Load Together

How do the three axes we have seen — settings, memory, permissions — operate at the same time in one session? Let's unfold it around a single line of input. The following is a cross-section of what actually happens when I type "Let's review skill cooldown balance."

settings.json Memory (JIT) Permissions UserPromptSubmit hook fires inject_memory.py runs (exit 0 guaranteed) manifest 17 atom score sort·regex top 3·6000 chars ceilings applied cooldown atom attached before input allow / deny checked per tool call reads auto · deletes need approval

Three lanes meet at one input. settings.json wakes the hook, the hook picks the memory and attaches it to the input, and when the response built that way calls a tool, permissions operate as the final gate. The user typed only one line — "check the cooldown" — yet three pieces of infrastructure take their turns out of sight. This is the internal structure of the moment a tool starts to feel like my tool.


1.3.6 First Setup Guide — Operational Within One Hour

We have seen the theory; now we move our hands. One hour after first installing Claude Code, you can lay the whole picture above onto your own PC. I split it into five segments.

0–10 minutes: install and confirm it runs. After installing, run Claude Code from a terminal. In some folder, ask "What's in this folder?" and confirm a response. First check that the tool is alive.

10–25 minutes: write three global memory files. For the auto-loaded global layer, three files are enough.

Copying my example verbatim is a fine way to start. Refine it as you operate.

25–40 minutes: basic settings.json setup. Set effortLevel to high, put in the starter permission set (reads and searches automatic; writes and deletes on approval), and take one backup (settings.json.bak_<date>). The backup is the single most important line in this segment.

40–55 minutes: your first five project atoms. Turn five things into atoms — the decisions you keep forgetting, the information you keep asking for. The folder is ~/.claude/projects/<project>/memory/. See Chapter 5 for the format. With five, you get results from global auto-load alone, without building a JIT manifest yet.

55–60 minutes: one test. Open a new session and throw one question from your own field. Check that global memory auto-loaded and that the response tone follows your collaboration rules.

That is the hour. The JIT manifest and the hook can wait until your atoms pass about 50, when auto-load starts to feel heavy — that is the moment to lay down the inject_memory.py from 1.3.2.


1.3.7 Common Mistakes and How to Avoid Them

The mistakes that recur in the early days group into five, and each stands on the same accident cause.

Mistake Accident cause How to avoid
Putting too much into global Every session grows heavy and wastes tokens Keep global under 5KB; move the details to project memory
Auto-allowing every permission Convenience hiding risk — the first seat of an accident Auto only reads/searches; writes/deletes on approval (with quarterly cleanup)
Editing settings without a backup A broken settings file keeps the tool itself from starting Save settings.json.bak_<date> automatically before every change
Piling atoms into the memory folder without limit Auto-load presses against the token ceiling Adopt a JIT manifest from around 50 atoms
Mixing team and personal settings in one file git conflicts, personal settings exposed Team in settings.json, yourself in settings.local.json

You do not need to dodge all five from day one. Global bloat and the missing backup deserve avoidance patterns within the first hour; for the other three, it is more natural to run for about a month and then fit the avoidance devices into the spots where your own accident probability runs highest.


1.3.8 Closing Part 1

1.1 was where we closed the distance to the tool, 1.2 where we grasped its minimum working mechanism, and 1.3 where we laid the first infrastructure of memory, permissions, and settings. These three chapters form the book's introduction. Finish this far and the basic skeleton for running the tool without stalling is in place.

The key point is that this skeleton is not a static configuration. The atoms in the manifest grow every week, the allow list lengthens along the traces of your work, and scores accumulate through verification. Infrastructure is not finished the moment it is laid — it grows together with its user on top of what was laid. The gap that opens between 17 on my personal PC and 304 at the company is the distance of that growth.

From Part 2 we enter information architecture proper. Chapter 4 covers YAML frontmatter, Chapter 5 atoms, Chapter 6 layers, and Chapter 7 ontology, in that order. The atom you saw in 1.3 as just one manifest entry becomes the protagonist of a whole chapter in 2.2. Only after the spine is set can the domain chapters find their own places on the same coordinates.


Key Takeaways

Next Chapter Preview


Try It Yourself

setup 1. Open ~/.claude/settings.json and, before changing anything, take a backup as settings.json.bak_<date> (the suffix is today's date). 2. Put Read(*), Grep(*), Bash(ls:*), and Bash(git status:*) into permissions.allow, and enter Bash(rm -rf:*) and Bash(git push --force:*) into permissions.deny. 3. (Once you have 50 or more atoms) Register python ~/.claude/hooks/inject_memory.py under hooks.UserPromptSubmit, and write the atom entries (name, path, regex, score) plus config.max_matches: 3 into _jit_manifest.json.

prompt - In a new session, ask a question that deliberately includes a keyword from one of the manifest's atoms. Example: "Review skill balance against the cooldown rules."

verify - Check that a signal like [memory injected: N atoms] appears right after the input. - See whether the intended atom is reflected in the response. - Deliberately put broken JSON into the manifest, confirm the prompt still goes through unblocked (the exit 0 guarantee), then restore it.

Solo Scale-Down - Start with no hook and no manifest. In a single global MEMORY.md, write just 3 lines of identity and 3 lines of collaboration rules, and auto-allow only Read(*) and Grep(*). When atoms grow familiar in your hands and approach 50, that is when you add the hook from 1.3.2. Infrastructure starts small and grows along your traces — you do not begin with 304.

Part 2 · Info Architecture

2.1 YAML Frontmatter — Every Document as Data

The night before a milestone build, teammate A, a systems designer, asked me on our team messenger: "How many documents touched the reward curve this week? And how far has review gotten?" I didn't know the answer. The documents were somewhere in the folders; who had last touched them, and which milestone they belonged to, lived scattered across individual memory and file-naming conventions. What we did that night was make one convention: six lines entered at the top of every document. From the next milestone on, those six lines made it possible to answer teammate A's question without anyone opening a folder.

A few lines of YAML written between --- markers at the very top of a document. This is called frontmatter. This convention tells both humans and machines, at the same time, what a document is — without reading a single character of the body. This chapter follows how that one line becomes the entry coordinate for the entire information architecture, with a script that actually runs.

One term up front. This book divides design documents into five Layers (Chapter 6 covers them in earnest): L0 = worldview and concept, L1 = system rules, L2 = content, L3 = data, L4 = implementation coordinates. Dependencies flowing top-down is the normal direction. The layer: 2 that appears in the section below is a coordinate declaration: "this document belongs to the content Layer."


2.1.1 Why "Documents as Data" Rather Than Just "Documents"

Traditional design documents have lived in Word, PowerPoint, and Google Docs. The body text is optimized for human reading. But the metadata — a document's kind, ownership, status, location — is either dissolved into the body or delegated to folder structures and file-naming conventions. So to learn "which milestone does this document belong to, who is responsible, when was it last reviewed," you have to open the body.

Two limitations compound here. First, a document does not declare its own identity. Its identity lives in people's memory and folder conventions, and those conventions corrode over time. Second, AI has no clues for inferring context. Tell Claude Code "review this document," and it reads the body start to finish, wasting tokens, with no idea where the scope of responsibility ends.

YAML frontmatter solves both at once. Put explicit metadata in the first lines of the document, and both humans and machines can identify it without opening the body. It is like a label on the front of a filing-cabinet drawer: you know what's inside without pulling it open. And this label is more than a classification tool. As we'll see later, the single layer field becomes the entry coordinate for procedural generation and automated review.


2.1.2 A Real Frontmatter — The First 14 Lines of One Document

Instead of an abstract example, here is the frontmatter that Project A's reward-curve document actually carries at its head (only IDs and real names are pseudonymized; the structure is exactly as it runs in production).

---
title: "Main Quest Chapter 12 Reward Curve"
layer: 2
status: review
owner: teammate_a
created: 2026-04-15
updated: 2026-05-20
related:
  - quest_main_chapter12
  - reward_curve_milestone_2
affects:
  - L3_BalanceSheet_v2
ip_check: passed
---

# Main Quest Chapter 12 Reward Curve

(body begins)

The point is the separation above and below ---. Above is data the parser reads; below is body text humans read. Markdown renderers usually hide frontmatter, so it does not interrupt reading. One file holds both data (frontmatter) and content (body), becoming a single source of truth.

Watch the two lines layer: 2 and affects: [L3_BalanceSheet_v2]. They declare: "this content (L2) document affects the balance sheet in the data Layer (L3)." From that alone, a tool can draw the L2→L3 dependency as a graph without reading the body. Conversely, if an L3 data document references an L1 system rule via depends_on — a reverse dependency pointing from bottom to top — that is a design smell. A tool detects that reverse reference automatically.

Why YAML is easier to write by hand than JSON is simple: indentation expresses structure, quotes are rarely needed, and # comments are allowed. It suits designers filling it in themselves.


2.1.3 Where the Standard Lives — _NAMING_FRONTMATTER_STANDARD

Fields can multiply without limit. The more they grow, the heavier the writing burden and the faster the standard collapses. So Project A runs two tiers: a minimal set of core fields common to every document, and domain-specific extension fields.

The common core is six fields.

Field Format Purpose
title string Human-readable title. May differ from the file name
layer 0–4 The Layer coordinate from Chapter 6
status draft / review / approved / archived Document state
owner username The person responsible (exactly one)
created YYYY-MM-DD Creation date
updated YYYY-MM-DD Last modified date

These six alone tell you a document's freshness, ownership, and position at a glance. Resist the urge to add more for the first month. As you operate, which fields you actually need reveals itself naturally.

Extension fields differ by domain. Systems design favors depends_on and affects; combat design, combat_phase and anim_target; narrative, world_region and chapter; balance, data_sheet and formula_id. These extension fields must not scatter freely, so a single standard document nails down each field's official name, allowed values, and examples. That document is _NAMING_FRONTMATTER_STANDARD.md. Adding a new field has to go through it. And the standard document itself is registered as an atom, managed in the same family as the rule that forces a Layer number prefix on document names (the docs_layer_numeric_prefix_naming atom).

An important shift happens here. If the standard is only a document humans read, humans will break it. Make the standard data that machines read, and machines enforce it. The next section is the actual code of that shift.


2.1.4 Worked Transcript — Enforcing the Standard in Code, and the Lesson from a datetime Bug

Now I had Claude Code build "a linter that checks whether every Markdown document in Project A follows the frontmatter standard." There were two core requirements. It had to catch the violations (missing required fields, non-standard status values, layer outside 0–4, documents in review untouched for more than 90 days), and it had to read allowed values from the standard document instead of hardcoding them. That separation is the point: change the standard, and the check criteria change without touching the code. (The full script and the steps to run it yourself are in the "Try It Yourself" section at the end of this chapter.)

Then came an incident. Claude's first version computed the date difference in the STALE check as today - fm["updated"], with a comment saying that if a file reads updated: 2026-05-20, PyYAML auto-parses it as a datetime.date. That is only half true. Run against the real documents, some files threw a traceback.

TypeError: unsupported operand type(s) for -: 'datetime.date' and 'str'

The cause was human hands. Some authors wrote updated: 2026-05-20 (parsed as a date); others wrote updated: "2026-05-20" with quotes (parsed as a string). Where the standard had not pinned down the date format, people's habits split — and Claude assumed only one side. I rejected the code and asked again: "normalize both notations safely to a date, and also handle the case where updated is missing." Claude inserted a helper that checks the input type and normalizes both to datetime.date (the fixed block is also in "Try It Yourself").

The real lesson was not the code bug. It was that people's habits split exactly where the standard had not pinned down the date format. So I added one line to _NAMING_FRONTMATTER_STANDARD.md: updated: YYYY-MM-DD (따옴표 없이) — that is, YYYY-MM-DD, without quotes. The linter, in the middle of doing its checks, ended up exposing a hole in the very standard it was checking against.

The fixed script's first output was not clean. I am leaving the messy result exactly as it came out.

[NO-FM]   manuscript/legacy/old_combat_notes.md
[MISSING] manuscript/system/quest_flag_table.md: layer
[STATUS]  manuscript/content/town_intro.md: WIP
[LAYER]   manuscript/balance/dps_v2.md: None
[STALE]   manuscript/system/inventory_rules.md: 134d

These five lines were the team's actual state in the early days of adoption. Old documents had no frontmatter at all (NO-FM), one document was missing layer, someone used the non-standard value status: WIP, one balance document left layer as None, and one system-rules document had been asleep in review for 134 days. A standard is never followed from day one. The linter merely surfaces that fact every morning.


2.1.5 From Frontmatter to Script — The Flow

Compressed into a single diagram, the worked transcript above looks like this. It shows how one line written by a human flows all the way into the machine's automated checks.

flowchart TD
    A["Author: creates a new document
template auto-inserts the 6 frontmatter fields"] --> B["Fills in only the blank values
title / layer / status / owner ..."] B --> C["Saved .md file
--- frontmatter --- + body"] C --> D["Linter script runs
rglob('*.md')"] D --> E["_NAMING_FRONTMATTER_STANDARD.md
allowed values read from here"] E --> F{"Checks
required fields / status / layer / stale"} F -->|"pass"| G["Feeds the relation graph
related·affects → L-coordinate dependencies"] F -->|"violation"| H["Daily report
auto-notifies the responsible owner"] H --> B G --> I["AI queries become possible
'Gather Layer 2 docs in review' → instant answer"] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef fail fill:#fee2e2,stroke:#dc2626,color:#7f1d1d; class A,B human class D,F code class C,E,G data class H fail class I ai

Two things matter. First, the standard (E) is separate from the script (D). Change the standard, and the check criteria change without touching the code. Second, a violation (H) is not a dead end but a loop back to the writing stage (B). It doesn't blame anyone; it routes the document back so its author fixes their own document.


2.1.6 A Production Case — Six Months on a Mid-Sized Team

On Project A, which I run as design director, we rolled out frontmatter to the entire design team (four to five people) about six months ago. Adoption did not happen in one stroke; it passed through four distinct stages.

The biggest pushback in week one was "you want me to write this by hand every time?" Memorizing and typing six lines for every new document is tedious. The fix was template auto-insertion. VSCode snippets, Obsidian templates, and the "New Document" button on our design portal all insert an empty YAML block automatically. Authors fill in only the blank values. The pushback disappeared within a week.

At one month, standard collisions erupted. With several people adding fields freely, owner, responsible, and author all appeared at once. Same concept, three spellings — both search and automation broke. The fix was to consolidate every field's official name, allowed values, and examples into the single document _NAMING_FRONTMATTER_STANDARD.md, and to make adding any new field go through that document as a rule. The standard stabilized within a month.

At three months, the linter from 2.1.4 came in. Even with a standard, people break it. So a consistency report began generating automatically every morning and landing in the shared channel of our team messenger. Each owner only has to look at their own documents. After automation, standard violations dropped noticeably (author's estimate, not a precise measurement — felt like half or less).

At six months, the combination with AI began to pay off. Once the standard stabilized, queries like these came back with instant answers.

In the end, frontmatter became a shared vocabulary between humans and AI. When a human writes it, AI understands it; when AI writes it, a human verifies it. Both look at the same keys. But remember the week-one pushback, the one-month collisions, the three-month linter, the six-month combination — six accumulated months produced this result. It did not appear in one stroke.


2.1.7 Common Mistakes and How to Avoid Them

The mistakes that recur in early adoption group into five. All of them stand on the same root — "places where the standard was left to human willpower alone."

Mistake Why It Goes Wrong How to Avoid It
Defining too many fields from the start Authors burn out filling in blanks; quality drops Start with the core six; after 1–2 months add only the ones you actually use
Field names keep changing (tagtagscategory) Old names linger in accumulated documents; search and automation break Pair every rename with a migration script. Auto-convert or warn when old names are found
Typing it by hand every time Typos, missing fields, and split date notations (the bug from 2.1.4) become routine Templates, snippets, "New Document" automation first. Human hands only for meaningful values
A standard left alone with no verification Even with a standard, nobody knows who broke it; natural corrosion sets in Linter + daily automated report, so violators fix their own documents
Forgetting the layer field Without a Layer coordinate, neither cross-discipline visibility nor review gates can form Make layer a required field. The linter detects omissions

You don't need to block all five from day one. Set up the avoidance patterns for #1 and #3 in week one; fold in #2, #4, and #5 as you operate, starting wherever your own team stumbles most often.


2.1.8 Starting Small — Making It Stick in Three Weeks

Adopting frontmatter is lighter work than you might expect. Three weeks is enough for it to settle into one team.

In week one, define the core six fields, build the template, and apply it to new documents only, keeping the writing burden minimal. In week two, apply it by hand to the top 20 most-viewed documents, and check in real use which fields are missing. In week three, turn on the linter and the daily report — from then on, the standard is maintained by the strength of tools, not human willpower.

Do not migrate the entire backlog of documents at once. Start with frequently viewed documents and with new documents. After about six months, nearly every document carries frontmatter. Even then, 100% is not the goal. Spending time migrating old documents nobody has ever opened is waste.


Try It Yourself

Run one full cycle yourself, at the smallest possible scale.

setup - Place two or three .md documents to be checked in a working folder. Deliberately remove layer from some, or insert a non-standard value like status: WIP. - Place a minimal standard document in the same folder. status: allowed = ["draft", "review", "approved", "archived"] updated: YYYY-MM-DD (without quotes)

prompt (enter into Claude Code)

Write a Python script that checks the YAML frontmatter of every .md under this folder. Catch missing required fields title·layer·status·owner, status values outside the allowed list (read it from the standard document), layer values that violate the 0–4 integer rule, and documents in review whose updated is more than 90 days old. Handle updated safely whether it comes as a string or as a date, and print violations per file.

verify - Run the script and confirm that every violation you planted gets caught. - Add WIP to the allowed list in the standard document and run again. Confirm that status: WIP now passes even though you didn't change a single line of code. That is the proof that the standard and the code are separated. - Include one document with quotes around updated and one without, and confirm the TypeError from 2.1.4 does not appear.

Reference: The Full Linter Script

This is the code Claude first produced in 2.1.4. The STALE check line (age = (today - fm["updated"]).days) still contains the datetime bug.

import sys, datetime, pathlib, re
import yaml  # PyYAML

ROOT = pathlib.Path("manuscript")
STANDARD = pathlib.Path("_NAMING_FRONTMATTER_STANDARD.md")
REQUIRED = ["title", "layer", "status", "owner"]

def load_allowed_status(standard_path):
    # Extract the allowed `status` values from the standard document
    text = standard_path.read_text(encoding="utf-8")
    m = re.search(r"status:\s*allowed\s*=\s*\[(.*?)\]", text)
    if not m:
        return ["draft", "review", "approved", "archived"]
    return [s.strip().strip('"').strip("'") for s in m.group(1).split(",")]

def parse_frontmatter(md_path):
    text = md_path.read_text(encoding="utf-8")
    if not text.startswith("---"):
        return None
    end = text.find("---", 3)
    block = text[3:end]
    return yaml.safe_load(block)

def main():
    allowed = load_allowed_status(STANDARD)
    today = datetime.date.today()
    violations = 0
    for md in ROOT.rglob("*.md"):
        fm = parse_frontmatter(md)
        if fm is None:
            print(f"[NO-FM]   {md}")
            violations += 1
            continue
        for field in REQUIRED:
            if field not in fm:
                print(f"[MISSING] {md}: {field}")
                violations += 1
        if fm.get("status") not in allowed:
            print(f"[STATUS]  {md}: {fm.get('status')}")
            violations += 1
        if not isinstance(fm.get("layer"), int) or not (0 <= fm.get("layer") <= 4):
            print(f"[LAYER]   {md}: {fm.get('layer')}")
            violations += 1
        if fm.get("status") == "review":
            age = (today - fm["updated"]).days   # ← this is where it breaks
            if age > 90:
                print(f"[STALE]   {md}: {age}d")
                violations += 1
    sys.exit(violations)

The core block that came back after my re-request. It normalizes updated safely whether it arrives as a string or as a date.

def as_date(v):
    if isinstance(v, datetime.date):
        return v
    if isinstance(v, str):
        return datetime.date.fromisoformat(v.strip())
    return None

# Replacement for the STALE check inside main()
if fm.get("status") == "review":
    upd = as_date(fm.get("updated"))
    if upd is None:
        print(f"[MISSING] {md}: updated")
        violations += 1
    elif (today - upd).days > 90:
        print(f"[STALE]   {md}: {(today - upd).days}d")
        violations += 1

Solo Scale-Down

You don't need a team. In a personal notes folder, cut the core fields down to three — title, status, updated — and have the linter catch only one thing: "documents whose status is review and whose updated is more than 30 days old." That alone surfaces, once a week, the documents you started reviewing and then forgot. The standard–template–check triangle works at solo scale exactly as it does for a team.


Key Takeaways

2.2 Per-Page Atoms — The Anatomy of One Decision per Document

During a new hire's first week, he asked me over chat: "Is the combat cooldown 0.6 seconds? Which document says so?" I answered, "It's in the skill system GDD (Game Design Document — the detailed spec)." He asked again: "Which section of that GDD? It's 220 lines — class design, then damage curves, then even how the UI displays it." I opened the file and found it for him. Line 137. Then he asked one last question: "But why 0.6 seconds? Was 0.5 not an option?" That answer was in no document at all. I remembered we had decided it in a meeting six months earlier, but the reason was buried somewhere in the meeting notes.

That five-minute conversation contains all three failures of a 220-line consolidated document. You cannot find the location (search failure), there is no reason (lost context), and a human has to mediate every time (no automation). Ask an AI the same question and things get worse: it reads all 220 lines and then mixes damage-curve talk — irrelevant to cooldowns — into its answer.

The prescription in this chapter is simple. One document holds one decision. A decision-sized document carved out by this principle is called an atom. Split the 220-line GDD and "the cooldown is 0.6 seconds" becomes one atom, and inside that atom the location, content, reason, exceptions, and relations all sit in one place. Instead of abstractions, this chapter dissects one real atom from end to end: how to name it, what frontmatter to fill in, how to make relations explicit, and how, as a result, the AI picks out exactly that one atom and nothing else.


2.2.1 Picking One Specimen — combat_cooldown_rule_v2

The specimen on the table is one atom actually in operation on Project A. Its name is combat_cooldown_rule_v2. The full file follows. It is not long — it holds only one decision. (The file appears verbatim in its original Korean; each key line is quoted in English as the dissection proceeds.)

---
name: combat_cooldown_rule_v2
title: "Combat Cooldown Rule — v2"
type: rule
layer: 1
status: approved
owner: Lee Minsoo
created: 2026-03-10
updated: 2026-05-12
applies_to: [skill_system, item_system]
---

# Combat Cooldown Rule v2

Why: To limit the number of skills usable at once, reducing the
burden of split-second decisions and preserving the meaning of combo inputs.

Rule: Every active skill has a global cooldown of 0.6 seconds + an
individual cooldown (defined per skill). While the global cooldown is
running, no active skill can be cast.

How to apply:
- Every new skill definition must specify an individual cooldown
- A cooldown column value of 0 in L3_SkillSheet violates this rule
- The build-stage consistency check detects violations automatically

Exceptions:
- Passive skills are exempt from this rule
- Ultimates use a separate gauge system (See: [[ultimate_gauge_system]])

Relations:
- affects: [[combat_dps_calculation_v3]], [[balance_curve_v3]]
- derives_from: [[principle_decision_load_reduction]]
- conflicts_with: [[skill_cancel_rule_legacy_v1]]
- requires: [[combat_input_buffer_system]], [[skill_system_v2]]
- is_a: rule
- part_of: combat_system_master

I will cut this one file into five parts: naming, frontmatter, single decision, relations, traceability. Only when all five are in place does an AI read this atom as "a unit that makes sense on its own."


2.2.2 Part ① Naming — The Name Itself Is a Coordinate

The file name is combat_cooldown_rule_v2. It is not a name picked on a whim; it has a three-segment structure — the domain prefix, the decision body, and the version.

combat_         cooldown_rule          _v2
└ prefix        └ decision body        └ version
  (which domain)  (what it decides)      (which revision)

The prefix combat_ is a coordinate that says "this is a decision in the combat domain." Project A's rule atoms are partitioned into domains by prefix: quest_ (quests), data_ (data operations), docs_ (document operations), meeting_ (meeting notes), portal_ (the design viewer). The prefix alone tells you whose area of responsibility a decision belongs to and where its influences come from.

When naming wavers, everything wavers. If the same decision exists twice, as skill-cooldown.md and as cooldown_skill_v2.md, search breaks, and so does the JIT matching that appears later in this chapter. So Project A pinned the naming rule itself down as an atom before anything else. That atom is atom_naming_convention_v1, and it mandates snake_case, a required prefix, and a version suffix. And the rule is enforced not by human willpower but by a linter. Commit a file name without a prefix and it gets caught at the build stage.

Beneath the naming lies a larger design that runs through this whole book. The frontmatter's layer: 1 is the second coordinate. If the prefix says "which domain," the Layer says "which layer of abstraction." Only when the two coordinates combine is the atom's position fixed as a single point on a plane. Here, Layer is just a coordinate (the detailed definition of layers 0–4 is in 2.3). The cooldown rule is "an input rule that governs generation," so it sits on Layer 1. There is even a separate rule that forces this Layer coordinate onto document names as a numeric prefix — docs_layer_numeric_prefix_naming. A single name carries two explicit coordinate axes.

The essence of this design is not a compulsion to tidy up. There is a line I repeated to the team: "The Layers were split for procedural generation in the first place." When every atom carries an explicit domain coordinate (prefix) and layer coordinate (Layer), an AI can later "take every Layer 1 combat rule as input and auto-generate Layer 2 content." The name is the addressing scheme for that automation.


2.2.3 Part ② Frontmatter — The Label Machines Read

The YAML block between the --- markers above the body is the frontmatter. It applies the standard from 2.1 to atoms as-is, and it is a label read not by people but by machines — the build script, the JIT hook, the relation map generator.

Field Value What machines do with it
name combat_cooldown_rule_v2 The unique ID that other atoms link to
type rule Per-category stats and filters (rule / concept / decision …)
layer 1 Per-Layer coloring and sorting; the reference axis for reverse-reference detection
status approved Of draft, approved, and archived, only approved gets into the build
applies_to [skill_system, item_system] Impact scope — the systems this rule touches
created/updated 2026-03-10 / 2026-05-12 Change tracking; the reference dates for auditing stale atoms

With these labels filled in, automated checks become possible. For example, if a system rule declared layer: 1 directly references a data atom (Layer 3) like [[L3_SkillSheet_row_0042]] in its body, that is a reverse reference (L3→L1) — an upper layer bound to a concrete value in a lower layer. Project A detects this pattern automatically at the build stage, because a rule should reference the format of the data, not a single row of it. Without that one layer line in the frontmatter, the check itself cannot exist.

Handling status: archived is also frontmatter's job. When a decision changes, the atom is not deleted; it receives status: archived plus an archived_at date. Build and JIT exclude archived atoms. The record stays, but it leaves active duty. Over six months of operation on Project A, the archive rate was about 15% (author's own measurement). If that ratio sits close to 0%, I read it as a signal that the archival workflow is not working.


2.2.4 Part ③ Single Decision — Does It Summarize in One Sentence?

The heart of the atom dissection is confirming that the body holds only one decision. The test is simple. Try to summarize the atom's decision in one sentence.

"Every active skill has a 0.6-second global cooldown."

One sentence does it. Pass. If the summary comes out as two sentences — "the cooldown is 0.6 seconds, and during a combo it is reduced by 50%" — that is two decisions. Split it into combat_cooldown_rule_v2 (base cooldown) and combat_combo_cooldown_reduction_v1 (combo reduction).

There are two more auxiliary tests for singleness.

The independent-retirement test. If you retire this one atom, does the system stay standing? Retire the cooldown rule and combat balance wobbles, but the system still runs. The unit is right. Conversely, if retiring it brings five other atoms down with it, those five are really five fragments of one decision. Merge them into a larger atom.

The single-reference test. Can someone elsewhere link just [[combat_cooldown_rule_v2]] and have it carry meaning? If it does, the unit is right. If referencing that one line forces a reader through several parts of the body, it has not been split enough.

A body that passes these tests naturally settles into five sections — Why, Rule, How, Exceptions, Relations. Above all, do not delete the Why. The answer to the new hire's final question from the opening — "Why 0.6 seconds?" — lives here: "to reduce the burden of split-second decisions and preserve the meaning of combo inputs." When someone proposes "let's cut it to 0.5 seconds" six months from now, this one line becomes the starting point of the debate. An atom whose Why is gone becomes a fossil nobody dares to touch.


2.2.5 Part ④ Relations — Arrows Make Impact Analysis Possible

The Relations section at the bottom of the atom is what makes this specimen a node in a graph rather than an isolated memo. The key point is that it does not just say "related documents" — it states the kind of each relation.

flowchart TD
    P["Principle: reduce decision load"] -->|derives_from| C["combat_cooldown_rule_v2"]
    C -->|affects| A1["combat_dps_calculation_v3"]
    C -->|affects| A2["balance_curve_v3"]
    C -->|requires| R1["combat_input_buffer"]
    C -->|requires| R2["skill_system_v2"]
    C -.->|conflicts_with| X["skill_cancel_legacy_v1
(conflict, pending retirement)"] classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef fail fill:#fee2e2,stroke:#dc2626,color:#7f1d1d; class P,C,A1,A2,R1,R2 data class X fail

Six kinds of relation each do a different job.

With a plain "Related: [Document A], [Document B]" link, a human has to work everything out one item at a time. With relation types entered as an enum, the machine works it out. "Show me everything affected if I change this atom" becomes an automatic query that follows affects, and "find every pair of rules that currently contradict each other" becomes an automatic check that scans conflicts_with. The full ontology design of these six enums comes in 2.4; 2.2 only notes that the atom standard applies that enum ahead of time.

The relation arrows are also input to the relation map generator. Project A's gen_relation_map.py reads every atom's frontmatter layer and Relations section and automatically draws an interactive relation map as HTML, colored by Layer. This is possible only because each atom carries a coordinate (Layer) and arrows (Relations).


2.2.6 Part ⑤ Traceability — The 30 Minutes One Atom Saved

An atom with all five parts in place is traceable. Who made this decision, when, and why, and what counts as a violation — all in one place. The value of traceability shows most clearly not in statistics but in the incidents it actually prevented.

Project A's meeting_image_caption_standard atom is a rule that every image attached to meeting notes must carry a caption stating which screen it shows, why it was attached, and what decision it relates to. Before this atom existed, a screenshot went into a set of meeting notes without a caption, and a week later a team member who saw it spent 30 minutes checking with the author to figure out "what screen is this?" After the atom existed, when the same omission recurred, the build-stage linter caught the caption-less image automatically. Five minutes to fix. 30 minutes became 5.

Another specimen, skill_listing_budget_wrapper_only_policy, is a rule that caps global slash command slots at 12: the actual skills live in a separate directory, and only the 12 wrappers are exposed globally. Before it was codified, global slash commands had ballooned to nearly 40 and chewed into the token budget at every session start. Since the atom was defined, an automatic cleanup tool trims the excess at the start of each session. The rule is enforced by tooling, not by human memory.

Project A has about 304 of these atoms accumulated (author's own measurement, at the six-month mark of operation). Looking only at the broad strokes of the distribution, recurrence-prevention rules (rule) take the largest share, followed by one-off decisions pinned down (decision), domain concepts (concept), and collaboration corrections (feedback). One atom saves minutes, but with 304 of them the cumulative savings cross into days. That is why I call atoms an "asset," not "housekeeping."


2.2.7 From Dissection to Automatic Injection — How JIT Actually Works

So far we have dissected one atom statically. Now watch it in motion. The JIT (Just-In-Time) hook from 1.3 picks only the atoms that match keywords in the input and injects them into context on the spot. The JIT manifest is a JSON file that maps each atom to its matching keywords and a score — here the regex catches both the Korean word for "cooldown" and its English forms.

{
  "name": "combat_cooldown_rule_v2",
  "path": "atoms/combat/combat_cooldown_rule_v2.md",
  "regex": "쿨다운|cooldown|글로벌 쿨다운|GCD",
  "score": 75
}

An actual injection flows like this.

flowchart TD
    A["User input:
What if we cut the skill cooldown to 0.5s?"] --> B["JIT hook:
scans the manifest regex"] B --> C{"Cooldown matched?"} C -->|"Yes score=75"| D["combat_cooldown_rule_v2
injected in full"] C -->|"No"| E["No injection"] D --> F["AI reads Why, Rule, and Exceptions
before responding"] F --> G["Answer: 0.6s rests on reducing
decision load. Cutting to 0.5s
requires re-reviewing the affects
targets: DPS calc and balance curve"] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; class A human class B,C code class D data class F,G ai

The last box is the point. The AI does not merely answer "it was 0.6 seconds." Having read the atom's Why, it gives the rationale; having read affects in Relations, it flags ahead of time what will shake if the value changes (the DPS calculation and the balance curve). All five dissected parts — split small, reasons written down, relations made explicit — come alive in the response.

This is where the one-decision-per-document principle reveals itself as the precondition for automation. If this atom had been the 220-line consolidated GDD, the moment the single word "cooldown" matched, class design, damage curves, and UI would all be injected wholesale, the token budget would shrink, and the AI would lose focus over which of five decisions to answer. The smaller and clearer the atom, the higher the JIT precision. Being finely split is not a virtue of tidiness — it is the precondition for automatic injection.

The score is the device that protects the context budget. When several atoms match one input, only the top N by score are injected (default: 3). The scoring criteria get settled through operation.


2.2.8 Personal Atoms and Team-Shared Atoms — Separating the Two Tiers

The dissected specimen combat_cooldown_rule_v2 is a team-shared atom that earned status: approved. Not every atom starts out in that seat. Project A splits atoms into two tiers.

The reason for the separation is psychological. Personal atoms must stay free, so that you can jot down an unverified hypothesis without pressure and retire it a week later. If everything were public to the team from the start, you would think "what if this turns out wrong" and never write it down at all. Conversely, team-shared atoms must be strict, so that everyone trusts and references them.

combat_cooldown_rule_v2, too, probably began as a one-line personal memo: "let's test a 0.6-second cooldown." After it was validated in an alpha build, it was promoted to team-shared as a change request, went through review by another designer, and became approved. This personal-to-team promotion flow is itself one axis of the self-improving loop by which an atom system gets smarter over time.


2.2.9 Five Common Mistakes

The mistakes that recur in the early days of atom operation come down to five. All of them grow from the same root: treating atoms as one-off memos instead of assets.

Mistake What breaks How to avoid it
Creating too many in the first week Unverified atoms pile up and operation collapses Start with one or two verified ones; let the count grow naturally
Never archiving Stale atoms keep matching in JIT and produce wrong answers Quarterly audit; status: archived + archived_at
Too abstract / too specific "Do good design" cannot be verified; a one-line stray thought is meaningless Aim for the level of "attack ranges are 0.5/1.5/3.0/5.0 only"
Inconsistent naming Search and JIT matching break wholesale Create the naming convention atom first and enforce it with a linter
Skipping the Why Over time it becomes a fossil nobody can touch Enforce the five sections: Why, Rule, How, Exceptions, Relations

You do not need to dodge all five perfectly in the first month. Mistakes 1 and 4 are solved together by a single naming convention atom, and 2, 3, and 5 fall into line naturally if you run one quarterly audit around the three-month mark.


2.2.10 On to the Next Chapter

In this chapter we cut one atom into five parts: the name (coordinates), the frontmatter (a machine label), the single decision (the one-sentence test), the relations (impact analysis), and traceability (the 30 minutes it saved). And we saw how all five come alive at once in JIT automatic injection.

Of the two coordinates written into the name, 2.2 only brushed past layer: 1. 2.3 takes that Layer head-on. Give every atom a Layer coordinate and, even across disciplines, you start to see where each other's outputs sit. Then 2.4 formalizes, as an ontology, the six relations whose enum names this chapter merely borrowed (affects, derives_from, conflicts_with, requires, is_a, part_of). In the backbone of the information architecture that runs YAML (2.1) → Atom (2.2) → Layer (2.3) → Ontology (2.4), this chapter was the second vertebra.


Key Takeaways


Try It Yourself — Build One Atom and Inject It via JIT

setup. Create an atoms/ directory in your working folder, and write the naming convention atom (atom_naming_convention_v1) before anything else. Three lines are enough: snake_case, required prefix, version suffix. If you use JIT, put an empty array in _jit_manifest.json.

prompt. Pick one decision you keep forgetting, and get an atom draft with the prompt below.

"Turn the following decision into the standard atom format. Decision: 'Active skills have a 0.6-second global cooldown.' Use five sections: Why, Rule, How to apply, Exceptions, Relations. In the frontmatter, include name (snake_case + prefix), type, layer, status: draft, owner, and created. At the end, check whether the decision summarizes in one sentence."

verify. Check the atom you receive in three ways. ① Does the decision summarize in one sentence? (If not, split it.) ② Is the Why non-empty? ③ Add one {"name", "path", "regex", "score"} line to the manifest, then type that regex keyword as an actual input — does the atom get injected? Pass all three and your first atom is done.


Solo Scale-Down

If you are a solo developer with no team, no linter, and no build pipeline, this entire chapter shrinks down to a single folder in your note app.

The point is not the tools but the habit of the five parts. The first 10 notes are the hardest; once past that hump, your hands make the next 100 on their own.

2.3 Layer Design — Abstracting Game Systems

It was the quarter our design disciplines were growing from three to eight. A combat designer locked a skill's range at 8m. The same week, a level designer locked a dungeon corridor width at 6m. Each decision was perfectly reasonable inside its own discipline. The problem surfaced three weeks later, in the build. An area-of-effect skill punched through the corridor walls, and enemies died where the player couldn't even see them. It was nobody's mistake. The two simply had no window into each other's decisions.

This chapter is about building that window: letting each discipline keep its own room while a single coordinate tells it what is happening next door. That coordinate system is what this book calls the Layer.


2.3.1 Siloing — The Enemy We Keep Meeting

Game design splits into finely specialized disciplines: systems, combat, narrative, content, level, balance, UX, QA. Each discipline has its own tools, deliverables, and meetings. As scale grows, each goes deeper into its own territory until nobody knows what the others are doing. This is called siloing.

The cost of siloing only reveals itself with time.

Lack of skill is not the cause. Everyone made reasonable decisions within their own discipline; there was simply no channel for noticing the decisions of the others. Patch the gap with meetings and meetings explode; patch it with group chat and the signal drowns in noise. The point is not that meetings and chat are worthless — it is to draw a clear boundary between what they can patch and what they cannot.

The solution is to keep each territory intact (preserving discipline specialization) while letting everyone see each other's flow (integrated visibility). The two demands look contradictory, but align them on a shared coordinate system and you achieve both at once. That coordinate system is the Layer. In office terms: everyone keeps their own desk while looking at the same wall clock and the same calendar.


2.3.2 Defining the Layer — A 5-Tier Abstraction

The Layer this book uses is a 5-tier abstraction, numbered 0 through 4. The higher you go, the more abstract and the more rarely things change; the lower you go, the more concrete and the more often they change.

Abstract · Invariant Concrete · Volatile L0 Vision · Core Values Procedural generation role: context anchor — invariant, the reference point injected on every call L1 Systems · World Skeleton Procedural generation role: generation input rules — rulebooks, relations, tags (constraints the generator follows) L2 Content · Flow Procedural generation role: where generated bodies accumulate — quests, progression, level curves L3 Implementation · Data Sheets Procedural generation role: values, IDs, relations — the inputs to simulation L4 Build · QA Artifacts Procedural generation role: verification gate — build results, bugs, play captures

The role each of the five tiers plays in a procedural generation and automation pipeline appears in the right-hand labels of the diagram above. That mapping is the spine of this chapter. If you see the Layer as nothing more than "well-organized folders," you are seeing only half of it. Each tier corresponds to exactly one stage of the generation pipeline (anchor → rules → body → values → gate) — the last of these being the verification gate, this book's term for what the wider industry would call a quality gate: a checkpoint where a human or a checker verifies output before it moves on.

Layer What it holds Change frequency
Layer 0 The core experience the game intends to give the player. Compressible into one sentence Very low (the project's entire lifetime)
Layer 1 The large structure of the game's systems and the skeleton of the world Low (per milestone)
Layer 2 Play flow, quest lines, progression stages, level curves Medium (per sprint)
Layer 3 Actual data values, parameters, formulas, variables High (daily)
Layer 4 Results confirmed in the build, bug reports, play footage Very high (real time)

These five tiers are not a game-only concept. The same spine transfers directly to general IT product development. If you have never made a game, use the job-translation table below to map each tier onto your own deliverables (left: the Layers of game design; right: what occupies the same slot in SaaS, apps, or internal systems).

Layer Game design General IT products Same question
L0 Core experience The core experience for the player (one sentence) Product vision — whose problem, solved how "Why are we building this?"
L1 System rules System structure, world skeleton Business and feature rules — domain rules, permission models, core workflows "What should work, and how?"
L2 Content Quest lines, progression stages, level curves Releases and roadmap — feature bundles, ship order, milestones "What ships when?"
L3 Data Data values, parameters, formulas Spec sheets — API specs, field definitions, config values, thresholds "What are the exact values and definitions?"
L4 Build/QA Build results, bugs, play footage Deployment and QA — deploy artifacts, bug reports, monitoring logs "Does what actually shipped run correctly?"

Read it exactly as you would for games. The higher tiers change rarely (product vision: once a quarter), the lower tiers change often (config values: daily). The silo accident we saw earlier — the skill range colliding with the corridor width — has exactly the same structure as the general-IT accident where "a backend field definition (L3) and a frontend screen rule (L1) drift apart and blow up right before launch." Only the discipline names differ; the spine is one.

These 5 tiers are not absolute. Depending on scale and domain, 4 tiers may be the right fit, or 6 may be needed. The point is not that the number is 5, but the act itself of defining the tiers explicitly.

A single deliverable can also straddle two Layers. A "skill system GDD (game design document — the detailed specification)" carries both system design (Layer 1) and concrete data (Layer 3) at once. In that case, either split the document, or keep its primary Layer at 1 and pull the data section out into a separate sheet — but whichever way you choose, state explicitly which Layer each part lives in.


2.3.3 The Meta-Principle — Specialization and Integration at Once

Disciplines spread horizontally; Layers stack vertically. One discipline's work spans multiple Layers. The matrix below shows the distribution's center of gravity for 11 disciplines (columns) × Layers 0–4 (rows), with cell darkness indicating weight. The darkest cell is that discipline's center-of-gravity Layer.

Systems Combat Narrative Content Level Balance UX/UI QA Character Art Live Ops L0 Vision L1 Systems L2 Content L3 Data L4 Build·QA Center of gravity Secondary Minimal/none

Read a column and you see which Layers one discipline spans; read a row and you see which disciplines gather on one Layer. In the L0 (vision) row, narrative and art direction are darkest — the two disciplines closest to the vision. In the L3 (data) row, systems, combat, level, balance, and character cluster darkly — a signal that this is where they collide with each other, in the data sheets.

Hold this distribution explicitly and another discipline instantly knows where to look — "I need to check combat's Layer 2." The silo walls don't come down; windows get cut into them.

Compress the whole matrix into one sentence: the vertical axis (Layers) was cut to automate generation, the horizontal axis (disciplines) was cut to preserve expertise — and the two meet in a single cell of the grid.


2.3.4 A Case Study — Real Measurements from an MMORPG Project

On Project A, the MMORPG where I serve as design director, the design team (4–5 people) has been running the Layer system for about six months (the full development team is mid-size, 10–50 people). Let's look at concrete examples.

First, the narrative 5 tiers. The narrative design folder itself is split by Layer.

NarrativeDocs/ Layer0_Vision/ The world's core message, 1.1~1.2 Layer1_World/ Regions, factions, era settings Layer2_StoryLine/ Main quest flow Layer3_DialogueSheet/ Actual dialogue and name data Layer4_BuildVO/ Voice-overs shipped in the build

When a narrative writer changes one branch of the main story at Layer 2, the change ripples into the Layer 3 dialogue sheets — and can have an irreversible effect on the already-recorded Layer 4 voice-overs. Because the Layers are explicit, the blast radius can be traced immediately.

A relation-map generator, gen_relation_map.py, runs alongside this. It analyzes foreign-key relations between data sheets, builds an interactive HTML relation map, and expresses each node's Layer as its color (red = L1 systems, yellow = L2 content, green = L3 data). Which Layer depends on which is visible at a glance. When a dependency flows backward — when L3 fires an arrow at L1 — it is almost always a design flaw.

The procedural level generation master document declares its Layer coordinates in its frontmatter.

---
title: Procedural Level Design Master v0.1
layer_inputs: [L1.World, L2.StoryLine]
layer_outputs: [L3.LevelData, L4.PlayCapture]
---

These two lines declare "this pipeline takes Layers 1 and 2 as input and produces Layers 3 and 4," and the impact-radius calculation on change becomes automatable. The L0 vision needs no explicit mention — it is always an input, because the vision anchor rides along with every generation.

There is also an atom rule that enforces a Layer prefix on document names. One of the team's shared atoms reads:

docs_layer_numeric_prefix_naming: Data sheet filenames must carry a Layer-number prefix (L1_, L2_, L3_). A sheet without the prefix gets a warning from the consistency check.

The simpler the rule, the more powerful it is. A plain name sort already groups files by Layer; AI tools know the Layer from the filename alone. And when a human forgets, the consistency check catches it.


2.3.5 Detecting Reverse References — A Worked Transcript

The previous section (2.3.4) said that "when L3 fires an arrow at L1, it is almost always a design flaw." What happens when you hand this detection to an AI instead of human eyes? Here is one breath of an actual run, transcribed without polish (only company-identifying information anonymized).

[Full Prompt]

Look at the two attached frontmatters and check whether the Layer dependency direction is normal.
References are only normal when they flow from higher numbers→lower numbers. Warn if reversed.

--- Document A (L1_combat_system.md) ---
title: Combat System Skeleton
layer: L1
refs: [L3_skill_damage_sheet, L2_combat_flow]

--- Document B (L3_skill_damage_sheet.md) ---
title: Skill Damage Sheet
layer: L3
refs: [L1_combat_system]

(The prompt hands Claude two frontmatter blocks and one rule — references may only flow from higher Layer numbers to lower ones — and asks it to check the dependency direction and warn on anything flowing backward.)

[Claude's Raw Output — Verbatim, No Summarizing]

I inspected both documents.

Document B (L3_skill_damage_sheet, L3):
  refs → [L1_combat_system]
  Direction: L3 → L1 (concrete → abstract)
  Verdict: Normal. A data sheet referencing the system
        skeleton it follows is the natural direction.

Document A (L1_combat_system, L1):
  refs → [L3_skill_damage_sheet, L2_combat_flow]
  - L1 → L2 : normal (design pointing at a downstream flow)
  - L1 → L3 : ⚠ reverse-direction warning
        The L1 system skeleton directly references the L3
        data sheet. This means "the design document depends
        on specific data values," creating a coupling that
        forces design-doc edits whenever the data changes.

Recommendation: Remove L3_skill_damage_sheet from document A's refs.
      If the system skeleton must know a concrete value, either
      lift that value into L1 as an abstract rule (e.g., "base
      damage = class coefficient × level"), or route via an L2 flow document.

Summary: 1 reverse reference detected (A: L1 → L3).

(Claude judges document B normal — a data sheet referencing the system skeleton it follows — and flags one reverse reference in document A, L1 → L3: a coupling that would force the design doc to change whenever the data changes. It recommends removing the ref, and either lifting the value into L1 as an abstract rule — e.g., "base damage = class coefficient × level" — or routing through an L2 flow document.)

[Human Review / Veto]

The AI's verdict is correct. But the recommendation to "route through an L2 flow document" was too much for this case. What L1 needed to know was the shape of the damage formula, not a specific cell in the sheet. So I adopted only the second recommendation (lifting it into an abstract rule) and vetoed the routing suggestion.

[Follow-Up Request]

Drop the routing suggestion, and show me only the edit diff for
L1_combat_system.md in the direction of "lift only the damage
formula's shape into L1 as an abstract rule." Clean up the refs too.

(The follow-up drops the routing suggestion and asks for only the diff to L1_combat_system.md that lifts the damage formula's shape into L1 as an abstract rule, with the refs cleaned up.)

This one cycle is the home ground of reverse-reference detection. The AI catches the direction violation (automated), a human trims the recommendation to its proper scope (review), and only the narrowed work is requested again (follow-up). On Project A, gen_relation_map.py does this at the graph level, and the portal_layer_change_impact_check atom fires at change-detection time, forcing an impact-radius check.

Had a person done this comparison by hand, opening both documents, matching the refs, and judging the direction takes several minutes. Once documents number in the hundreds, it becomes practically impossible. Reverse references always slip in quietly, one or two at a time, and only blow up in a build much later.


2.3.6 Layer Decomposition = The Precondition for Procedural Generation and Automation

The surface purpose of Layer integration is dissolving silos and unifying the language of collaboration (2.3.1–2.3.5). The essential purpose runs one level deeper. Once Layer decomposition takes root, the preconditions for procedural generation and automation are in place.

The case studies in the previous two sections were the stage where humans decide and AI assists with verification and injection. The next stage is where AI drafts a discipline's mass production itself as candidates, and humans adopt them. Layer decomposition is the precondition for that move, for three reasons. (1) AI candidate generation must be able to specify "generate what, on which Layer." (2) Automated consistency checking only works when the dependency direction between Layers is standardized (the reverse-reference detection of 2.3.5). (3) Automatic change-impact calculation requires a coordinate for which Layer a change happened on. All three converge on one statement: without Layer decomposition, automation itself is blocked. At the end of the hand that drew these coordinates, procedural generation was waiting from the very start.

There is no need to go deep before you have seen the discipline parts, so I will only sketch the two stages of application. Conservative application: humans decide, and AI automatically assists with consistency checks, change-impact calculation, and just-in-time (JIT) injection — the case studies of 2.3.4 and 2.3.5 live here. The tooling cost is small, the cumulative effect shows around month six of operation, and most mid-size (10–50 person) teams can reach it. Progressive application goes one step further: AI generates a discipline's mass production itself as candidates (narrative personas, PCG rulebooks, procedural levels, balance-change candidates, art assets, and so on), and humans decide only "which candidate to adopt." Each discipline's form and tool maturity are covered in that discipline's part.

Progressive application needs three ingredients common to every discipline: (1) Layer separation and labeling infrastructure (frontmatter, atoms, filename prefixes); (2) a candidate generation–evaluation cycle (N AI candidates → automated evaluation → ranking with a reasoned report); (3) a human review gate (only adopted results move on to the next Layer). At no point, however, does the deterministic core (simulation, physics, legal constraints) leave the hands of humans and deterministic code, and every review concludes in a reversible stage — before entering irreversible stages (recording, casting, live exposure, and so on). This reversible/irreversible boundary is a principle shared across all disciplines.

One last note on timing. Conservative application was partially possible even in the 2010s, but progressive application was blocked by three limits: the expressiveness of AI candidate generation, natural-language interpretation in automated evaluation, and the human review burden. With the advance of LLMs, all three entered practical territory, and progressive application came down from a vision on paper to working practice. Here lies the meta-message that runs through this entire book: the progress of AI has raised the feasibility of procedural generation and automation.


2.3.7 Discipline Coordinates — Where This Book's Discipline Parts Live

Each discipline part of this book states in its introduction which Layers that discipline mainly occupies, and uses Layer coordinates freely within its chapters. Here is the summary in advance (it transposes the centers of gravity from the 2.3.3 matrix into a table).

Discipline Primary Layers Notes
Systems design L1–L3 Broad, from design skeletons to data sheets
Combat design L1–L3, some L4 Combo skeletons to damage sheets; build measurements
Narrative design L0–L4 Runs the full 5-tier structure as folders
Content design Centered on L2 Progression flow, quest lines
Level design L2–L3 Includes the procedural generation pipeline
Balance design Centered on L3, measures at L4 Data values, curves, verification measurements
UX/UI design L1–L3 Interaction skeletons to screen data
QA design Centered on L4, verifies L0–L3 Verifies that every Layer made it into the build
Characters, pets, mounts L1–L3 Systems, world, data
Art direction L0–L1 + L4 artifacts Vision and world guides + build review
Live ops L2–L4 Operations cycles, real-time data

Every discipline touches other Layers too, but knowing the center of gravity reveals the collaboration channels. Balance (L3) and live ops (L2–L4) meet at L3, so they must always work closely together; the two disciplines nearest to vision (L0) are narrative and art direction. These adjacencies surface naturally on the coordinate system.


2.3.8 Start Small, Grow Big

Try to adopt the Layer system perfectly from day one and you will never start. Bringing it in gradually is the right answer.

flowchart LR
    S1["Stage 1
Single-discipline adoption
(recommended: narrative)"] --> S2["Stage 2
Adjacent-discipline expansion
(narrative + content)"] S2 --> S3["Stage 3
All-discipline standardization
(filename prefix, consistency checks)"] S3 --> S4["Stage 4
AI tool integration
(JIT, change impact, retrospective classification)"] S1 -.->|"Small (~10) teams"| E1["Enough up to here"] S3 -.->|"Mid-size (10~50) teams"| E2["Recommended up to here"] S4 -.->|"Large (100+) teams"| E3["Go this far for full effect"] classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; class S1,S2,S3,S4 human class E1,E2,E3 pass

Each stage takes at least a month, sometimes a quarter. Push too hard and people burn out. Pacing it so the operating burden never exceeds the adoption value — that is the director's job.

Small teams (\~10) need stages 1–2, mid-size teams (10–50) stage 3, and large teams (100+) must go to stage 4 for the full effect. That does not mean small teams can't use it. Only the depth differs; the core value already begins at stage 1.


2.3.9 Conclusion — The Spine of the Entire Book

The Layer is not a mere folder-organizing technique. It is the meta-principle that binds specialized game design into a single coordinate system AI can reason over — and, beyond that, the shared precondition for procedural generation and automation in every discipline.

Every remaining part of this book presupposes this chapter. The discipline parts state each discipline's Layer coordinates in their introductions; the process parts cover operating systems that cut across Layers; and the operations part covers the self-improving cycle of the Layer system itself.

The next chapter (game ontology and knowledge graphs) adds semantic relations on top of the Layers. If the Layer is the coordinate, the ontology is the semantic arrow drawn on those coordinates. Only when the two combine does AI autonomously infer that "this document affects that one."

One thing must be made clear. No automation in this chapter made a decision in anyone's place. In the reverse-reference detection, the machine merely laid out violation candidates; what to accept, and how far, was chosen by a human hand. The Layer is a coordinate system that helps people make better decisions faster — not a device for handing decisions off.


Key Takeaways

Key Atoms of This Chapter (Reference)

Next Chapter Preview


Try It Yourself

setup — Pick one discipline folder (recommended: narrative) and split it into five subfolders, from Layer0_Vision/ to Layer4_BuildVO/. Move your existing files into their Layers. For files that are hard to place, use "how often this document changes" as the criterion (the more often it changes, the lower the Layer).

prompt — Pick two data sheet frontmatters, paste in the full prompt from 2.3.5 as-is, and have the AI judge the Layer dependency direction. One precise rule is all it needs: "references must flow only from concrete to abstract (higher numbers to lower numbers)."

verify — When the AI catches a reverse reference, do not accept its recommendation wholesale; trim it to the right scope yourself (the "Human Review / Veto" of 2.3.5). Re-request a diff for only the direction you adopted. When a name sort shows your files grouped by Layer from top to bottom, the prefix rule has taken hold.

Solo Scale-Down

The Layer works even when you work alone. With no team there is no cross-discipline silo, but there is a cross-time silo: the you of three weeks ago and the you of today forget each other's decisions. Split your folders into nothing more than Layers 0–4 — one vision page, a few system skeleton pages, a progression flow, data sheets, build notes — and you instantly find which slot your past self put things in. Attach a single line for the AI, "currently working at L2," and it stops dragging in material from unrelated Layers. Trimming down to 4 or 3 tiers is fine. The number is not the point; the act of making the tiers explicit is.

2.4 Ontology and the Wikilink Graph — Verifying the Semantic Arrows

Monday morning, a change request comes in. Team member A on the combat team posts one line in the team messenger: "Changing the global cooldown from 0.5s to 0.3s. Does this affect anything?" Normally this is where a 30-minute meeting starts. The owner of the damage formula raises a hand, the owner of the combo-cancel rules cuts in, and someone asks, "Doesn't this hit the boss patterns too?" Nobody holds the whole picture in their head, so the meeting fills up with people digging through their memories.

This time is different. One second after the request goes up, a bot posts an automatic comment: "Changing this atom affects 4 atoms: skill_dps_calculation, combat_combo_cancel_v3, refgame_boss_pattern_phase2, balance_curve_v3. Owners: team member B, team member A, team member C." No meeting was held. Four people each checked their own atom, and that was the end of it. (We build this bot ourselves later — 2.4.3.)

This comment is not magic. In 2.3 we gave every atom a Layer coordinate, and in this chapter we add semantic arrows on top of it — which decision affects which. A coordinate only says "something is here." Relations like "this affects that," "that must exist first," "these two must never be enabled together" are arrows drawn on top of the coordinates. This chapter covers how to notate those arrows and how to catch broken ones automatically.

Terminology Notes - Ontology: a system that explicitly defines concepts and the relations between them. This book uses a lightweight version simplified down to 6–12 relations. - wikilink: an inter-document link in the [[atom_name]] format. The notation is borrowed from Obsidian, Roam, and similar tools. - Backlink: the list of "atoms that point at this atom" — the reverse direction of a forward reference. - Orphan node: an atom referenced from nowhere. A signal that it is a deprecation candidate. - Broken link: a wikilink pointing at an atom that does not exist. The residue of typos and renames.


2.4.1 Relations Are Arrows — Why Wikilinks Alone Are Not Enough

In 2.1 we attached metadata with YAML frontmatter, and we scattered wikilinks through atom bodies. That alone already connects the documents into a mesh. The problem is that the connections say nothing about what they mean.

This decision stands on [[skill_cooldown_rule_v2]].

This one line — "this decision stands on [[skill_cooldown_rule_v2]]" — says only that the rule is mentioned. Why is it mentioned? Does this decision need that rule (requires), was it derived from it (derives_from), or does it conflict with it (conflicts_with)? A person reading the sentence knows; a machine does not. Ask the AI "does enabling this decision break anything?" and links without meaning cannot answer.

So we put a relation type on each wikilink. The relations actually used in game design work are surprisingly few. The following six cover more than 90% of cases.

affects influences derives_from derived from requires must exist first conflicts_with mutually exclusive is_a special case part_of part of e.g. combat_combo_cancel_v3 —[affects]→ skill_dps_calculation combat_combo_cancel_v3 —[derives_from]→ vision_taste_focused_combat combat_combo_cancel_v3 —[requires]→ combat_input_buffer_system

The atom that freezes these six as an enum is ontology_relation_enum_v1. Adding a new relation type means going through change-request review. Even as the set grows, 10–12 is the right ceiling, and starting with just three — affects, derives_from, requires — is plenty. The place to write relations is the atom's YAML frontmatter.

---
name: combat_combo_cancel_v3
layer: 1
affects: [skill_dps_calculation, refgame_boss_pattern_phase2]
derives_from: [vision_taste_focused_combat]
requires: [combat_input_buffer_system, skill_cooldown_rule_v2]
conflicts_with: [skill_cancel_rule_legacy_v1]
---

A person writes only the forward direction, one line. The reverse direction ("who affects me?") is computed by a tool scanning the whole folder. Write affects on both ends and the synchronization burden doubles; the moment you fix one side and forget the other, the graph starts lying. One direction by hand, the other direction by tool — that is the principle running through this entire chapter.


2.4.2 Worked Transcript — Verifying Wikilink Backlinks and Graph Integrity in One Pass

From here on I reproduce one real session in full. The situation: while replacing the combo-cancel rule v2 with v3, we decided to deprecate skill_cancel_rule_legacy_v1. But before deprecating it, we have to check whether any other atom points at this atom (backlinks) and whether the deprecation leaves broken links or conflicts behind (graph integrity). Grep by hand and you will miss things. I hand it to Claude.

The Prompt (in Full)

The attached atom folder (decisions/) contains atoms, and their frontmatter
should carry relations (affects/derives_from/requires/conflicts_with/is_a/part_of).
The bodies contain wikilinks too. I want to deprecate skill_cancel_rule_legacy_v1
— before deprecating, run 3 checks: 1) backlinks 2) broken links
3) conflicts. Don't guess; actually read the files and answer. If you can't find something, say so.
// backlinks = both YAML relations + body wikilinks / broken links = links that go stale or dangling on deprecation
// conflicts = whether any conflicts_with counterpart is still alive with status: approved

(In English: "The attached atom folder, decisions/, contains atoms; their frontmatter carries relations (affects/derives_from/requires/conflicts_with/is_a/part_of) and their bodies contain wikilinks. I want to deprecate skill_cancel_rule_legacy_v1 — before that, run 3 checks: 1) backlinks 2) broken links 3) conflicts. Don't guess; actually read the files and answer. If you can't find something, say so. // backlinks = both YAML relations and body wikilinks / broken links = links that go stale or dangling on deprecation / conflicts = whether any conflicts_with counterpart is still alive with status: approved")

Claude's Output (Raw — Not Summarized, Not Cleaned Up)

(Claude's output below is translated from the Korean original; identifiers, counts, and structure are unchanged.)

I read the 47 files in the decisions/ folder. Here are the results of the 3 checks for skill_cancel_rule_legacy_v1.

1. Backlinks (atoms that point at this atom)

Atoms pointing at it via YAML relations: - combat_combo_cancel_v3conflicts_with: [skill_cancel_rule_legacy_v1] - skill_cancel_rule_v2derives_from: [skill_cancel_rule_legacy_v1]

Atoms pointing at it via body wikilinks: - combat_combo_cancel_v3 body: "solves the input-buffer problem of the legacy [[skill_cancel_rule_legacy_v1]]" - qa_combat_regression_checklist body: "regression item 3 was written against [[skill_cancel_rule_legacy_v1]]"

So 4 atoms reference this atom. (2 via YAML + 2 via body wikilinks; since combat_combo_cancel_v3 appears on both sides, the actual atom count is 3.)

2. Broken links (what breaks on deprecation)

Deprecation by itself does not "break" any wikilink — the file stays, and only its status becomes deprecated. However: - qa_combat_regression_checklist defines a regression item against the deprecated atom, so it becomes invalid in substance. That is not a broken link, but it is a stale link. Needs review. - If you go as far as deleting the file, all 4 references above become broken links.

⚠️ One ambiguous point: the body of balance_curve_v3 contains a link [[skill_cancel_rule_legacy]] (no _v1 at the end). Whether this is a typo meant to point at _v1, or intentionally refers to a separate atom, I cannot tell from the files alone. There is no file named skill_cancel_rule_legacy in the folder. So this is most likely an already-broken link. It needs confirmation.

3. Conflict residue (is the conflicts_with counterpart alive?)

Atoms tied to skill_cancel_rule_legacy_v1 by conflicts_with: just one, combat_combo_cancel_v3. - status of combat_combo_cancel_v3: approved (alive)

→ The new rule that conflicted with the deprecation target is in normal approved status. This is fine. It is the natural replacement — the old version retires and the new one survives. Not a conflict-residue warning.


Summary: deprecation is safe, but handle two things first. (a) The stale link in qa_combat_regression_checklist — update regression item 3. (b) [[skill_cancel_rule_legacy]] in balance_curve_v3 — suspected typo. May already be a broken link.

Human Verification and Rejection

Here I did not take the output at face value. (a) is correct — the regression checklist is mine to update. But I stopped at (b). Claude called the [[skill_cancel_rule_legacy]] in balance_curve_v3 a "suspected typo," but I had to look at the folder myself to see whether it really was a broken link. An AI's "I cannot tell" is an honest signal, not a conclusion.

I opened the folder and checked: a file named skill_cancel_rule_legacy really did not exist. But the git log showed a commit from 6 months earlier that had done one thing — renamed skill_cancel_rule_legacyskill_cancel_rule_legacy_v1. The body link in balance_curve_v3 never got updated then and slipped through. So this was not a typo; it was a broken link orphaned by a rename. Claude's "suspected typo" guess was only half right. The cause was not a typo but an unpropagated rename. Had I not caught that difference, I would have fixed only the regression checklist and the broken link would have stayed buried.

The Follow-Up Request (Reflecting the Confirmed Facts)

Confirmed. [[skill_cancel_rule_legacy]] was not a typo — it was a broken link
from a rename 6 months ago (_legacy → _legacy_v1) whose body link never got fixed.

So while we're at it, find every "wikilink pointing at a nonexistent atom"
across the whole folder. By body wikilinks, everything whose target file is missing from the folder.
A table of each file and the broken links inside it.

(In English: "Confirmed. [[skill_cancel_rule_legacy]] was not a typo — it was a broken link from a rename 6 months ago (_legacy → _legacy_v1) whose body link never got fixed. So while we're at it, find every wikilink in the whole folder that points at a nonexistent atom — by body wikilinks, everything whose target file is missing from the folder. A table of each file and the broken links inside it.")

I collected every [[...]] body wikilink across the 47 files and checked them against the folder's file list. Broken links whose target file does not exist: 3.

File containing the link Broken wikilink Suspected cause
balance_curve_v3 [[skill_cancel_rule_legacy]] rename not propagated (→ _v1)
quest_reward_table_v2 [[concept_reward_currency]] target atom never written (planned, but no file)
refgame_boss_pattern_phase2 [[skill_global_cooldown]] rename not propagated (→ skill_global_cooldown_rule)

Of the 3, the 2 rename cases can be fixed by correcting the target atom name to _v1 / _rule. In the 1 remaining case (concept_reward_currency), the target atom itself does not exist — you either write it or remove the link. I did not fix anything automatically — confirm which you intended, then proceed.

All three were real. Grepping by hand, I would almost certainly have missed the broken link in quest_reward_table_v2. That link was an intentional forward reference — pointing ahead at an atom not yet written — and after 6 months in which nobody wrote that atom, it had become, in effect, a dead promise.

What this session shows is simple. Backlink detection and broken-link detection are work the AI is strong at — reading an entire folder and cross-checking it — while cause determination and intent confirmation belong to the human. The AI goes as far as "this is broken here"; the human goes as far as "why it broke and how to fix it."


2.4.3 What a Graph Makes Visible — Turning Verification Visual

You could rerun the previous section's checks as a prompt every time, but freeze the same checks into code and they become visible at a glance on a graph. On Project A, a graph tool extending the gen_relation_map.py introduced in 2.3 runs as an R&D effort. The core is to read the folder's atoms, build a networkx directed graph, and lay four check functions on top.

import networkx as nx

# build_graph(folder): reads the atom folder and builds a DiGraph
#   from nodes (=atoms) and YAML relation edges. (Full source in "Try It Yourself")

def find_cycles(G):                      # circular dependencies
    return list(nx.simple_cycles(G))

def find_orphans(G):                     # inbound 0 = orphan candidate
    return [n for n in G.nodes if G.in_degree(n) == 0]

The heart of it is two lines. simple_cycles catches circular dependencies (A requires B requires C requires A), and in_degree(n) == 0 catches orphan nodes — no need to write your own DFS. The remaining two functions are one-liners of the same grain. find_broken_wikilinks collects the body [[...]] links with a regex and picks out the ones missing from the node list, and backlinks fall out of walking the graph in reverse (full source in "Try It Yourself"). For visualization, color nodes by Layer and edges by relation type, and draw heavily referenced nodes (those with many inbound edges) larger so the hubs stand out. Like folders in a cabinet sorted with color labels, the patterns surface in your field of view before you go hunting for them.

Below is a graph of the actual relations among the atoms covered in the 2.4.2 session. The arrow direction means "the source atom asserts a relation toward the target atom."

graph LR
    combo[combat_combo_cancel_v3
L1·approved] legacy[skill_cancel_rule_legacy_v1
L1·deprecated] v2[skill_cancel_rule_v2
L1] dps[skill_dps_calculation
L3] vision[vision_taste_focused_combat
L0] buffer[combat_input_buffer_system
L1] boss[refgame_boss_pattern_phase2
L2] qa[qa_combat_regression_checklist
L4] broken[skill_cancel_rule_legacy
does not exist] combo -->|affects| dps combo -->|affects| boss combo -->|derives_from| vision combo -->|requires| buffer combo -->|conflicts_with| legacy v2 -->|derives_from| legacy qa -.stale.-> legacy balance[balance_curve_v3] -.broken.-> broken classDef dep fill:#eee,stroke:#999,stroke-dasharray:4 classDef miss fill:#fff,stroke:#cc2222,stroke-dasharray:4 class legacy dep class broken miss

The two dashed edges are the problems caught by human verification in 2.4.2. qa → legacy is the stale link against a deprecated atom; balance_curve_v3 → skill_cancel_rule_legacy is the broken link pointing at a node that does not exist. Drawn as a graph, those two dashed lines stand out among the solid ones. Run everything as text alone and they would have stayed buried somewhere across 47 files, never to be seen.

Four rules run automatically at the verification gate (Layer 4) — this book's term for what the industry calls a quality gate, the stage where automated checks pass or block a change.

Once these four rules harden into code, there is no need to write a prompt each time as in 2.4.2. The moment a change request goes up, the bot rebuilds the graph and posts the list of affected atoms plus any broken links, cycles, and conflicts as an automatic comment. The "4 atoms will be affected" comment at the top of this chapter is exactly this — the bot promised earlier.


2.4.4 Why Layers — Coordinates Drawn for Procedural Generation

Here we need to pause on why 2.3 and 2.4 belong together. Layer coordinates and relation arrows were not adopted separately; they are two faces of the same purpose.

On the surface, Layers unify the language of collaboration — call one thing "an L1 system decision" and another "L3 data," and people in different disciplines share the same coordinates. But the essential purpose lies elsewhere. Layers were coordinates drawn for procedural generation.

L0 vision is the context anchor — immutable, injected into the AI every time. L1 systems are the input rules of generation — rulebooks, relations, and tags live here. L2 content is where the generated body accumulates; L3 data is the simulation input of values, IDs, and relations; L4 build/QA is the verification gate. The relation arrows operate on these coordinates as constraints on generation. When the AI creates new content, a requires arrow becomes the precondition "this must exist first," and a conflicts_with arrow becomes the prohibition "these must not be enabled together."

Disciplines keep specializing (combat, quests, and economy each hold their own expertise), yet every artifact carries a Layer coordinate, so they stay aware of one another. Specialization and integration hold simultaneously on one coordinate system. Once this graph has grown enough, the AI generates candidates while reading the relation arrows as automatic constraints, and the human only confirms violations at the review gate. The worked transcript in 2.4.2 is the miniature of that — the AI read the graph and found the constraint violations (broken links, conflicts), and the human ruled at the gate.


2.4.5 Domain Concepts Are Nodes Too — A Small Shared Dictionary

The nodes at the start and end of relation arrows are not only decision atoms. Atoms that define domain concepts like "skill," "quest," and "reward" are first-class citizens of the graph. concept_skill_definition_v1 looks like this.

# Skill
Definition: A unit action a character triggers in combat. Comprises input, cooldown, cost, and effects.
Required Properties: input / cooldown / cost / effects
Subtypes: active_skill (is_a) / passive_skill (is_a) / ultimate_skill (is_a)
Not a skill: auto attack [[concept_auto_attack]] / transformation [[concept_transformation]]

(The definition reads: a skill is a unit action a character triggers in combat, comprising input, cooldown, cost, and effects; auto attacks and transformations are explicitly not skills.)

Write boss_skill is_a skill and the parent concept's rules are inherited automatically. Project A has about 19 of these concept-definition atoms, and every decision atom references them. When the vocabulary is unified into one dictionary, the translation burden between meetings, documents, and code disappears. It is keeping one shared dictionary instead of a different one on every desk. The concept_reward_currency caught as a broken link in the worked transcript earlier was exactly that — "an entry promised for the dictionary but never written."


2.4.6 Keep It Light — The Trap of Academic Ontologies

Read this far and the question comes up: "isn't this a formal ontology like OWL/RDF?" No. And it was deliberately built not to be one.

Academic ontologies (OWL, RDF, SKOS) are powerful, but they come with dozens to hundreds of relation types, require dedicated inference engines, and need an ontology specialist on staff to operate. In domains where precise inference is life-or-death — search engines, medicine, law — they are essential. But the inference game design actually needs is only "change impact scope," "prerequisite dependencies," and "conflict detection," and all of it ends at the level of simple traversal — BFS/DFS, walking the graph one step at a time. The previous section catching cycles with one line of networkx is the proof. A formal inference engine is overkill.

There is one criterion. Can a game designer operate it by hand? Six YAML enums, networkx processing, pyvis or D3.js visualization. The moment you cross that line, the tool stops being a tool and becomes one more burden. Lightness is not a compromise; it is design intent.

Five recurring mistakes undermine this. All of them grow from the same root: treating the ontology as an enforced standard.

Mistake How to avoid it
Defining too many relation types up front Start with three (affects, derives_from, requires); add only when needed
Forcing relations onto every decision Accept atoms with no relations as normal — empty relations pollute the graph
Writing affects in both directions One direction by hand; compute the reverse automatically with the tool
Fixating on OWL/RDF Stay at an operable level (YAML + enum)
Running text-only with no visualization Ship even a plain HTML view from day one — like the dashed lines in 2.4.3, what you cannot see, you do not fix

Mistakes 1, 2, and 3 settle into a pattern within the first month of adoption; review 4 and 5 at the three-month retrospective and they fall into line naturally.


2.4.7 Starting Small, and the Next Chapter

The first month runs at a loss. The burden of writing relations piles up while no visible payoff arrives. So start small. In the first week, use only the three relations affects, derives_from, requires; in weeks 2–4, apply them to 20 core atoms and watch the graph grow. Build one HTML graph view at the level of the previous section in the first month, and from the second month the visualization starts showing its value; by the third month, automated verification (cycles, conflicts, orphans, broken links) is directly cutting meeting time. Enduring that one month of deficit is the fork between adoption succeeding and failing.

Chapter 8, Wikilink, digs into the [[...]] notation this chapter treated as arrows, at the operational level — how to use the backlink panel daily, how to update links in bulk on a rename (how to prevent the unpropagated rename of 2.4.2 in the first place), and how to fold the graph view of tools like Obsidian into day-to-day practice. YAML (Chapter 4) → Atom (Chapter 5) → Layer (Chapter 6) → Ontology (Chapter 7) → Wikilink (Chapter 8) — the completed pentagon of the information architecture.


Try It Yourself

setup. Pick the folder where your atoms live (e.g., decisions/) and get ready to write relation keys in each atom's YAML. Start with just three: affects, derives_from, requires. Unify body links in the [[atom_name]] format.

prompt. Before any deprecation or rename, throw this at the AI.

Read the markdown atoms in this folder. I want to deprecate/change [target_atom].
(1) backlinks: every atom that points at this atom via YAML relations or body wikilinks
(2) broken links: links that break or go stale on change/deletion
(3) conflict residue: any conflicts_with counterpart with status: approved
Check these. Don't guess — actually read the files, and if you can't find something, say so.

(In English: "Read the markdown atoms in this folder. I want to deprecate/change [target_atom]. Check (1) backlinks: every atom that points at it via YAML relations or body wikilinks, (2) broken links: links that break or go stale on change/deletion, (3) conflict residue: any conflicts_with counterpart with status: approved. Don't guess — actually read the files, and if you can't find something, say so.")

verify. Open every spot the AI marked "suspected typo" or "cannot tell" yourself. The cause of a broken link — typo, unpropagated rename, or never written — is for the human to rule on, with the git log and the actual folder. Don't let the AI auto-fix; confirm the intent, then fix it yourself.

The Graph Tool in Full (Excerpted in 2.4.3)

The body text (2.4.3) showed only the two core functions. The full version that builds the graph from the whole folder and runs the four checks follows.

import networkx as nx
import re, yaml, glob, os

REL_TYPES = ["affects", "derives_from", "requires",
             "conflicts_with", "is_a", "part_of"]
WIKILINK = re.compile(r"\[\[([a-zA-Z0-9_]+)\]\]")

def build_graph(folder):
    G = nx.DiGraph()
    files = {}
    for path in glob.glob(os.path.join(folder, "*.md")):
        name = os.path.splitext(os.path.basename(path))[0]
        text = open(path, encoding="utf-8").read()
        fm = yaml.safe_load(text.split("---")[1]) or {}
        files[name] = fm
        G.add_node(name, layer=fm.get("layer"), status=fm.get("status"))
    # YAML relation edges
    for name, fm in files.items():
        for rel in REL_TYPES:
            for tgt in (fm.get(rel) or []):
                G.add_edge(name, tgt, type=rel)
    return G, files

def find_broken_wikilinks(folder, known_nodes):
    broken = []
    for path in glob.glob(os.path.join(folder, "*.md")):
        text = open(path, encoding="utf-8").read()
        for m in WIKILINK.findall(text):
            if m not in known_nodes:
                broken.append((os.path.basename(path), m))
    return broken

def find_orphans(G):
    # inbound 0 and no part_of/is_a parent either
    return [n for n in G.nodes if G.in_degree(n) == 0]

def find_cycles(G):
    return list(nx.simple_cycles(G))

Solo Scale-Down

You don't need the tooling. One atom folder and Claude are enough. When you write a new decision, add just two YAML lines — requires and affects — and whenever a deprecation or rename comes up, run the prompt above once. Graph visualization is a problem for later. The habit of sweeping for broken links and conflicts right before a change — that single pass is what prevents the biggest deficit in solo operation.


Key Takeaways

Part 3 · System Design

3.1 The Systems Designer's Work and Layer Coordinates

Thursday, 4:50 p.m. The skill sheet the balance designer had been filling in had just landed. 312 skills. Each skill has a column called effect_id where an effect number goes, and that number points to a row in a separate effect sheet. The two have to line up for the game to run. If they don't, the client calls an empty effect or dies quietly.

I used to check this by hand. One cell in the skill sheet, jump to the effect sheet, confirm the number, jump back. 312 times. Two hours at best. In the last 50 rows, eyes glazing over, I always missed one or two — and those one or two blew up in the QA build.

This chapter is about where those two hours went, and about which coordinate that consistency check occupies on the systems designer's map of work. If you don't fix the coordinates first, you will forever be deciding by gut feel where to plug in AI.


3.1.1 A Systems Designer Makes Four Things

The systems designer is the person who travels the widest range between the abstract and the concrete. They take the fog called vision and haul it all the way down to the hard numbers in the last cell of a data sheet. Four kinds of deliverables come out of that journey.

(1) Translating vision into structure. When the director says "action combat where every hit lands with real impact," the systems designer turns that into a skeleton of skills, combos, cancels, and hitstop. "Player agency over growth" becomes class, skill tree, and equipment systems. It is the first moment fog turns into structure.

(2) Specifying the interfaces between systems. Combat, movement, inventory, shops, quests, and guilds all run at once. Does opening the inventory mid-combat grant invincibility? What if a PvP request arrives in the middle of an enhancement? The answers to these cases add up to the feel of a "well-made" game. Every spot where an answer is missing is a spot where players feel friction.

(3) Owning the data sheets and their schemas. Coefficients for 312 skills, effects for hundreds of items, behaviors for dozens of monster types. The values they either fill in directly or hand off to the balance and content disciplines. But the column definitions (the schema) of the sheets stay in the systems designer's hands. It is the job of building labeled drawers. If the drawers are sloppy, everyone fills them differently and consistency breaks.

(4) Designing behavior logic. Character and monster AI comes out in forms like finite state machines (FSM), Behavior Trees (hereafter BT), decision tables, and procedural rules. This material goes to the programmers and becomes code.

The key point is that all four meet on one person's desk. That is why "what do I spend today on" becomes the systems designer's biggest operational decision.


3.1.2 System Deliverables Have Layer Coordinates

In 2.3 we placed every game-production artifact on a coordinate axis running from L0 (vision) to L4 (build). Now we plot the four deliverables from 3.1.1 directly onto that axis. Few disciplines scatter their deliverables across as many Layers as systems design does.

The following map shows, on a single page, where each deliverable lives on the Layers and whom you meet at each coordinate.

L0 L1 L2 L3 L4 Vision — the systems designer only receives it System skeleton: defining classes, combat, inventory, guilds Interfaces: inter-system interaction rules and priorities Schema + data: sheet column definitions, some values Build — QA verifies the intent is reflected ↔ Director / narrative ↔ Art direction ↔ Other systems designers ↔ Balance / content ↔ QA The segment the systems designer builds directly (L1→L3)

This map says two things. First, the systems designer is responsible for the long distance of receiving L0 and making it reach L4. Second, the segment they build with their own hands is L1–L3, and the collaboration partner changes at each of those three bands. Every band switch changes the collaboration language, so if you are not conscious of the coordinates, meetings keep spinning their wheels.

This does not mean one person touches all of L1–L3, though. On a large team, the L1–L2 owner and the L3 owner split. On a small team, one person covers it all. The coordinates are a map for dividing roles, not an order to dump everything on one person.


3.1.3 Once the Coordinates Are Set, the Slots for AI Become Visible

The map is drawn; now we color it in. Which coordinates get the biggest return on AI? Not a blanket "automate everything" — you choose by the nature of each coordinate.

flowchart TD
    L1["L1 system skeleton
(class count, combat model)"] L2["L2 interfaces
(interaction rules)"] L3["L3 schema + data
(312-row sheet)"] L1 -->|"tied to game identity
humans decide"| H1["AI = assists with variations/verification of human decisions"] L2 -->|"case explosion
wide impact"| H2["AI = automated extraction of change impact scope"] L3 -->|"structured, repetitive
consistency checks"| H3["AI = owns generation, verification, conversion"] H1 --> R["humans make the core decisions,
AI handles detail, consistency, repetition"] H2 --> R H3 --> R style L3 fill:#fff8e1,stroke:#f9a825 style H3 fill:#e8f5e9,stroke:#2e7d32 style R fill:#e3f2fd,stroke:#1565c0

The key is that the lower the coordinate, the larger the share AI can own outright. L1's "how many classes should we have" is game identity, so a human must hold it. At the other end, L3's "do all 312 rows of foreign keys line up" is structured and repetitive, so AI should take it whole. L2 sits in between — the human makes the decision, while AI backs it by extracting the impact scope: "if we change this rule, how far does it shake?"

This diagram explains why every exercise from 3.1.4 onward starts near L3. It is the spot with the biggest payoff and the smallest risk. A malfunctioning schema tool causes no accident, the relation map only draws pictures, and the consistency check can be vetoed by a human.


3.1.4 Worked Transcript: Handing the Consistency Check to AI at Coordinate L3

That is it for theory. Now back to that Thursday at the start of 3.1. We have AI check whether the effect_id of all 312 skills matches the effect sheet. No summarizing — I transcribe the exchange exactly as it happened.

Data entry follows the schema-first principle: a human first defines the spec sheet, the Enum and .proto come out of it (via a VBA — Excel's macro language — export), and the csv data is filled in last. So the verification prompt also starts with "read the schema first." And it nails down that the source of truth is not the schema document but the actually exported JSON (json_over_schema_doc_as_source_of_truth). Documents can lie; exported JSON cannot.

The Prompt (in Full)

You are our project's data consistency checker. The input is two exported JSON files. 1) skill.json — each skill object has skill_id, name, and effect_id fields. 2) effect.json — each effect object has effect_id and effect_type.

Check the following. (A) Does every effect_id in skill.json exist in effect.json (foreign key integrity)? (B) Orphan effects — present in effect.json but referenced by no skill. (C) Rows in skill.json where effect_id is empty or 0.

Rules: do not guess. If a field is not in the JSON, report it as "field missing." Output in three sections (A)(B)(C); for each violation, one line with the skill_id (or effect_id) and the reason. If a section has no violations, write only "no violations" in it.

Claude's Raw Output (Untouched, Unpolished)

I have not received the two input JSON files. No file contents are attached to the current message. Please resend one of the following: - The actual contents of skill.json and effect.json (pasted or attached as files) - Or the paths to the two files

Once I have the contents, I will check in (A)(B)(C) order. I will not fabricate results without the files.

Stopping here was, if anything, the signal of trustworthiness. If a plausible report had come back on empty input, that very neatness would be the red flag. The one that asks "give me the files first" instead of conjuring inspection results out of thin air is the one you can trust. I actually pasted the two JSONs in and asked again (only excerpts appear here, for length).

// skill.json (excerpt)
[
  {"skill_id": 1001, "name": "Heavy Strike",   "effect_id": 5001},
  {"skill_id": 1002, "name": "Spinning Slash", "effect_id": 5002},
  {"skill_id": 1003, "name": "Shield Raise",   "effect_id": 0},
  {"skill_id": 1004, "name": "Poison Coat",    "effect_id": 5099}
]
// effect.json (excerpt)
[
  {"effect_id": 5001, "effect_type": "DAMAGE"},
  {"effect_id": 5002, "effect_type": "DAMAGE"},
  {"effect_id": 5003, "effect_type": "BUFF"}
]

Claude's Second Output (on the Real Data)

(A) Foreign key integrity — 1 violation - skill_id: 1004effect_id: 5099 does not exist in effect.json (dangling reference)

(B) Orphan effects — 1 violation - effect_id: 5003 (BUFF) → referenced by no skill

(C) Empty/0 effect_id — 1 violation - skill_id: 1003effect_id: 0. Could be an intended "no effect" — needs confirmation against the schema whether 0 is a valid value.

Human Verification and Veto

Two of the three were clear bugs. 5099 was a number the balance designer had written into the skill sheet before adding the row to the effect sheet — the classic out-of-order edit. Caught.

But (C) I vetoed. In our project, effect_id: 0 is a valid agreed-upon value meaning "passive slot, no effect." The AI does not know that, so it raised "needs confirmation" — a sound judgment, but not something to count as a violation. So I moved to the next step: codify this convention as an atom, so that from the next check onward the AI knows 0 is a valid value.

The Follow-Up Request (Injecting Domain Knowledge)

Additional rule: effect_id == 0 is a valid agreed-upon value in our project, meaning "passive, no effect." Therefore effect_id == 0 is not a (C) violation. Report only empty strings, null, and negative numbers as (C) violations.

(C) Empty/0 effect_id — no violations (effect_id: 0 treated as a valid agreed-upon value. No empty/null/negative values.)

This one-line convention is exactly how domain knowledge accumulates. Systems design decisions are relatively permanent, so once you pin one down, the same check runs on the same standard a year later. One more decision card stacked in the cabinet.

In this transcript the human did only three things — (1) specified the input order: read the schema first; (2) confirmed that 5099 was a real bug; (3) knew that 0 was valid, and vetoed and corrected the AI's call. The two hours of jumping cell by cell through the other 312 rows are gone. What got automated was the labor of jumping and cross-checking; the three lines of judgment that remain are the core.


3.1.5 Accumulating Assets: Hardening the Coordinates into Code

You could order up the transcript's check by hand every time, but hardening work that repeats at L3 into tools is systems design orthodoxy. I will cite the two I actually operate — not abstract "Project A tools" but things that genuinely run on my desk.

gen_relation_map.py analyzes a sheet's column names and values, auto-detects foreign key relationships, and produces an interactive HTML relation map. Where 3.1.4 had a human drawing the arrow skill.effect_id → effect.effect_id in their head, this script draws that arrow as a picture across the entire sheet set. Spots where a dependency runs backward (the danger of L3 data reverse-referencing the L1 skeleton) jump out of the picture instantly.

The schema-doc skill parses the xlsm's $스키마 (schema) sheet and auto-generates a Markdown schema document. The very schema behind 3.1.4 (C)'s question — "needs confirmation against the schema whether 0 is a valid value" — becomes something a person can read fresh without digging through other files. When the sheet changes, the document follows, which shrinks the chronic disease of documents drifting away from the actual data.

Restating the two tools' positions in coordinates: schema-doc guards the column definitions at L3, and gen_relation_map.py guards the relationships between L2 and L3. AI-assisted prompts (verification like 3.1.4) run on top of them. The three are not separate acts — they cover different heights of the same coordinate axis.

You will walk through these tools hands-on in 3.2, 3.3, and 3.4. 3.1 was the map for deciding where to plug things in; the next three chapters do the plugging.


3.1.6 Gradual Adoption: Start from the Low-Risk Coordinates

Switch on all three tools at once, and the operating burden arrives before the benefit. In my experience the safe order starts from the coordinates where the risk is small (the lower ones).

Timing (suggested) Adoption Coordinate What happens if it goes wrong
Month 1 Schema first (3.2) L3 A document misses one update, nothing more
Months 2–3 Relation map visualization (3.3) L2–L3 A diagram is inaccurate, nothing more
Months 3–6 AI-assisted prompts (3.4) L1–L3 Everything passes verification, so a human can veto

The timeline is not an absolute standard. Depending on team size and existing infrastructure, it could take twice as long or finish in half the time (author's estimate, unverified). What does not change is the order. By putting the riskiest decision support last, you touch the most sensitive spot only after the first two tools have already built the team's verification habit.


3.1.7 Measurement — Honestly

I record the numbers without polish. What follows was observed on the design team (4–5 people; the full dev team is mid-sized at 10–50; about 6 months of operation) of the MMORPG project I run as director ("Project A" below). It is not precise automated instrumentation but the author's observation based on work logs and retrospective notes — read it for direction and rough proportions only.

The point is that the saved time is not time spent not making the game. It flows back into the deep decisions AI cannot take — things like the L1 skeleton. Cut the labor and spend it on judgment — that is this chapter's one-line recommendation.


Try It Yourself: One Consistency Check at Coordinate L3

setup. Export two sheets (say, skills and effects) to csv. If you can, convert them to JSON as well (the principle that the export artifact, not the document, is the source of truth). Pick one foreign key pair between the two sheets (e.g., skill.effect_id → effect.effect_id).

prompt. Use the full prompt from 3.1.4 as is. Do not drop the three key lines — (1) "read the schema/structure first," (2) "do not guess; report what is missing as missing," (3) "if there are no violations, write only 'no violations.'"

verify. Go through the AI's violation list line by line, as a human. Fix the real bugs; veto the false positives caused by domain conventions (e.g., 0 = 효과 없음 — "0 = no effect") and add that convention to the prompt (or an atom). If the same false positive disappears from the next check, you have banked one more asset.

Solo Scale-Down

If you are a solo developer with no team and no sheets, two Google Sheets tabs are enough. One tab is "skills," the other "effects." Link the two with a single effect_id column. Download the tabs as csv and paste them into the 3.1.4 prompt — dangling references and orphan effects get caught just the same on a 30-row sheet as on 312 rows. Only the scale differs; the coordinate is the same. Start at L3, and once it feels natural, climb one band at a time to relation maps and impact scopes.


Key Takeaways

3.2 Schema First — The $schema Sheet Comes Before the Data

Monday morning. A new designer had filled in 120 rows of the skill sheet; we built them to csv, and 28 red lines appeared in the client log. class_id references entry 47, but the class sheet has no 47. In the element column, someone wrote Fire, someone else wrote fire, and one row reads 화염 — "flame," written out in Korean. Tracing 28 red lines by hand, one at a time, eats half an afternoon.

The cause of this incident is not that the data was wrong. It is that nobody wrote down, before the data was created, the rules the data had to follow. When the rules live only in someone's head, the rules change the moment the person changes. This chapter covers a workflow that creates the rules — the schema — before the data. And it puts the enforcement of those rules into the hands of tools, as documents, not into human hands.


Terminology Notes - Schema: the column definitions of a data sheet — name, type, range, foreign keys, description. - $스키마: a dedicated column-definition sheet kept inside the Excel data file (xlsm). The name is Korean for "$schema." It holds only the rules for the columns, not data rows. - FK (foreign key): a column that references the PK (primary key) of another sheet — the way class_id points at a row in the Class sheet. - proto: a Protocol Buffers definition (.proto). The data-structure and Enum contract shared by client and server. - Single source of truth: the operating principle of managing each piece of information in exactly one place, so that everyone looks there.


3.2.1 The Input Order Is the Schema

Read "schema first" as merely "define the columns in advance" and you have caught only half of it. The heart of it is the order in which things get entered. The order in which the data-entering hand moves decides whether consistency holds or collapses.

The input order this book recommends is a four-box pipeline.

flowchart LR
    A["$스키마 sheet
(column rule definitions)"] --> B["Enum / *.proto
(code contract via VBA Export)"] B --> C["csv data
(rows filled within the rules)"] A -.->|schema-doc| D["Schema doc
(auto-generated .md)"] C -.->|gen_relation_map.py| E["FK relation map
(auto-generated HTML)"] D -.-> F(("AI / humans
read the same definitions")) E -.-> F classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; class A,B,C,D,E data

The solid arrows flowing left to right are the enforced input order. Define the $스키마 first, pull Enums and proto out of it with a VBA (Excel's macro language) Export, and fill in csv data only within that contract. The dotted lines are the artifacts that derive automatically from that input — the schema document (schema-doc) and the FK relation map (gen_relation_map.py) — and humans and AI read the same definitions through these derivatives.

As long as this order is enforced, most of the 28 red lines from the opening of this chapter are closed off before the data is filled in. If the fact that element is one of the four values fire/ice/lightning/none is pinned down as a proto Enum, both Fire and 화염 get caught at input time. If the fact that class_id references the Class sheet's PK is stated in the $스키마, the missing 47 is caught first by a check, not by the build.

Flip the order — fill the data first and tidy up the schema later — and the schema becomes after-the-fact cleanup. When you rework column rules on top of 1000 already-accumulated rows, the rules end up following the data, and at that moment the source of truth stands upside down.


3.2.2 Worked Transcript — From $스키마 to csv in One Pass

Instead of explaining in the abstract, I will run one sheet through the whole pipeline, start to finish. Say we are building a new skill sheet. Below is the complete record of the work, done with AI assistance. I am not summarizing; the places where things went wrong and where a human rejected the output are left in as they happened.

Step 1 — A Human Writes the $스키마 by Hand First

No tools, no AI yet. A human defines the column rules directly. This one step is never delegated.

# Skill sheet $스키마 (written by a human)
Column         Type       Range/Constraint     FK                  Description
skill_id       int        1000~9999            (PK)                Unique skill ID
name           string     1~30 chars           -                   Display name
class_id       int(FK)    -                    Class.class_id      Usable classes
cooldown       float      0.0~120.0            -                   Cooldown (sec)
damage_coef    float      0.0~10.0             -                   Attack power multiplier
element        enum       fire/ice/lightning/none  -               Element
is_passive     bool       true/false           -                   Passive flag

Step 2 — Ask the AI for a proto Enum Draft (Full Prompt)

The task: carry the element enum and the FK structure of the $스키마 over into a proto contract. This is format conversion, the kind of work you can hand to an AI.

Prompt (full text): Below is the column spec of a game data sheet. Convert this spec into a Protocol Buffers (proto3) message and Enum. Rules: (1) define each enum-typed column as a separate enum, and its first value must start with _UNSPECIFIED = 0. (2) For FK columns, leave the reference target as a comment. (3) Range constraints cannot be expressed in proto, so leave them as comments only. Do not add any column that is not in the spec.

skill_id int PK / name string / class_id int FK→Class.class_id /
cooldown float 0~120 / damage_coef float 0~10 /
element enum{fire,ice,lightning,none} / is_passive bool

Step 3 — Claude's Raw Output (As Is, Untouched)

syntax = "proto3";

// Skill sheet data contract
// Note: range constraints (cooldown 0~120, damage_coef 0~10)
//       cannot be enforced in proto. Check them in the linter.

enum SkillElement {
  SKILL_ELEMENT_UNSPECIFIED = 0;
  SKILL_ELEMENT_FIRE = 1;
  SKILL_ELEMENT_ICE = 2;
  SKILL_ELEMENT_LIGHTNING = 3;
  SKILL_ELEMENT_NONE = 4;   // ← human review marker
}

message Skill {
  int32 skill_id = 1;       // PK
  string name = 2;
  int32 class_id = 3;       // FK → Class.class_id
  float cooldown = 4;
  float damage_coef = 5;
  SkillElement element = 6;
  bool is_passive = 7;
}

Step 4 — Human Verification and Rejection

Here the human rejects one thing. The AI placed none at the end of the enum, as SKILL_ELEMENT_NONE = 4 — but in the $스키마, none is the default state meaning "no element." The point of contention is proto's value 0. In proto3, an enum's 0 is the slot that gets filled automatically when no value is provided, so the question is whether to keep 0 reserved as _UNSPECIFIED (value never entered) or to fill it with a semantic value like none (deliberately element-less). Merge the two into one slot, and a row whose cell was left blank (a mistake) and a row that deliberately chose no element (an intent) both arrive as 0, forever indistinguishable. The AI did keep _UNSPECIFIED = 0 and none separate, but it pushed none out to slot 4, leaving the most common default state far from searches and debugging.

The decisions the human made: - Keep _UNSPECIFIED = 0 (for detecting omissions). - Leave none where it is, but add to the csv authoring rules: "no-element must be written explicitly as none; blank cells are forbidden." Blank = 0 = UNSPECIFIED is treated as an input error. - Record this decision as one added line in the element row's description in the $스키마: "No element must also be explicit (none). A blank cell is an error."

The AI's output was not used as is. The format was accepted; the semantic boundary was drawn by a human.

Step 5 — Export Enum/proto from the Sheet with the VBA Export Button

The verified proto definition is emitted, consistently, by the Excel Export macro (a button on the $스키마 sheet). No human retypes it in an IDE. The sheet is the source of truth, so the proto comes from the sheet. This order is an extension of the "output is canonical" principle (json_over_schema_doc_as_source_of_truth) covered in 3.3 — the document does not explain the code; the sheet gives birth to the code.

Step 6 — Only Then Fill In the csv Data (AI-Assisted, Retry Included)

Prompt (full text): Create 5 csv rows that satisfy the Skill proto and $스키마 above. Constraints: class_id must be one of [1,2,3] (the PKs that currently exist in the Class sheet); damage_coef must be 0.0 for passives (is_passive=true); element must be the literal strings fire/ice/lightning/none; no blank cells.

One row in Claude's first output was off.

skill_id,name,class_id,cooldown,damage_coef,element,is_passive
1001,화염베기,1,3.5,2.4,fire,false
1002,빙결의손,2,8.0,3.1,ice,false
1003,체력회복,1,0.0,1.2,none,true     ← rejected: passive but damage_coef≠0
1004,번개창,3,5.0,2.8,lightning,false
1005,방어태세,2,0.0,0.0,none,true

Row 1003 violates the rule (is_passive=true requires damage_coef=0.0). The human rejected it and asked again.

Retry (full text): Row 1003 violates the rules. is_passive=true but damage_coef is 1.2. Passives must be 0.0. Fix only row 1003 and send it again.

Claude's revised output: 1003,체력회복,1,0.0,0.0,none,true

That the AI did not get everything right on the first try is not a flaw; it is simply something that happens. What matters is that because the schema was already in place, that one deviant row could be spotted by eye and corrected with a single line. Without the schema, row 1003 would have been discovered after the build, as an in-game bug where a passive skill deals damage.

The lesson of this whole transcript is simple. When the input order is fixed as $스키마 → proto → csv, the AI fills in the format fast and the human reviews only meaning and violations. When the order collapses, the human carries everything, from format all the way to meaning.


3.2.3 schema-doc — So No One Copies the Schema by Hand

Keeping the $스키마 inside Excel is convenient for designers, but to AI, git, and external tools it is a closed room. So we run a tool that automatically converts the $스키마 into Markdown. The slash skill schema-doc does this job.

It works in four steps.

  1. Parse the $스키마 sheet of the Excel file (xlsm) (python-calamine, Rust-accelerated)
  2. Extract the five elements of each column definition
  3. Convert them into a Markdown table
  4. Generate <sheet-name>_schema.md (the sheet's name plus _schema.md) in the same folder

The key point is that no human writes the schema twice. Define it once in Excel, and the Markdown is produced by the tool. The two cannot diverge. The trap covered in 3.3 — "make the schema document canonical and it drifts from the actual output" — is avoided here by flipping it: Excel is canonical, the document is derived.

What schema-doc generates (for the Skill sheet from the transcript above):

# Skill sheet schema  (auto-generated — do not edit directly)

| Column | Type | Range/Constraint | FK | Description |
|---|---|---|---|---|
| skill_id | int | 1000~9999 | (PK) | Unique skill ID |
| name | string | 1~30 chars | - | Display name |
| class_id | int(FK) | - | Class.class_id | Usable classes |
| cooldown | float | 0.0~120.0 | - | Cooldown (sec) |
| damage_coef | float | 0.0~10.0 | - | Attack power multiplier |
| element | enum | fire/ice/lightning/none | - | Element. No element must also be explicit (none); a blank cell is an error |
| is_passive | bool | true/false | - | Passive flag. If true, damage_coef=0 |

_source: Skill.xlsm / generated by schema-doc_

Notice how the boundaries the human drew in steps 4 and 6 of 3.2.2 carried straight into the description cells of element and is_passive. A human wrote one line in the $스키마, and now the document, the proto, and the validation all share the same rule. This is what a single source of truth looks like when it actually works.

Once the schema lands as Markdown, it is put to use immediately in three places.


3.2.4 gen_relation_map.py — A Graph That Shows Whether FKs Are Alive

If the schema is the rules inside a sheet, FKs are the rules between sheets. The definition that class_id references the Class sheet is written in the $스키마, but whether that reference is actually alive at this moment requires a separate check.

gen_relation_map.py automatically detects the FK relationships across the data sheets and draws them as an interactive HTML relation map. When arrows like Skill's class_id→Class and Item's set_id→ItemSet gather on one screen, an "FK whose target has disappeared" stands out as a broken arrow. An incident like the missing 47 from this chapter's opening becomes visible while the data is being filled in — as a severed line on the relation map, not as a red line in the build log.

The worked usage and visualization of this tool are covered in earnest in 3.3. From this chapter, remember one thing: if the $스키마 does not state the FKs, there is no graph for the relation map or the consistency check to draw. Stating FKs is not optional; it is a precondition of schema first.


3.2.5 The Five-Step Schema-First Workflow

Generalize the transcript of 3.2.2 and you get five steps. Separate each step's owner from its output, and it becomes clear what a human keeps in hand and what gets handed to tools.

Schema-First 5 Steps — Owner × Output Step Owner Output 1. Schema design Human $스키마 5 elements · FK definitions 2. Auto documentation schema-doc Schema .md 3. Contract extraction VBA Export Enum / *.proto 4. Data draft AI + Human csv rows (reject violations · retry) 5. Consistency · impact Linter / Relation map Violation report · FK graph Blue = human decision / Green = tool automation / Yellow = AI draft + human review

You do not need all five steps in the first month. Running just steps 1 and 2 (schema design + automatic documentation) captures half the value. Attach steps 3–5 gradually, once the operation has settled in. Enforce all five from day one, and the authoring burden will stall the rollout before it takes root.


3.2.6 What I Measured on Project A

On an MMORPG project I run as design director (hereafter "Project A"), I ran this workflow for about six months. Of the figures below, data-sheet column consistency and new-sheet drafting time are actual measurements aggregated from tool logs and work records; FK breakage frequency is an author's estimate (unverified), back-calculated from build-failure issues.

Item Before After Basis
Column name consistency about 60% about 95% measured via schema-doc comparison
FK breakage frequency 2–3 per week 1 or fewer per month back-calculated from build issues (author's estimate)
New sheet drafting time 4–8 hours 1–2 hours measured from work records
New designer understanding a sheet 3 meetings 1 document read + 1 meeting onboarding cases (directional only)

The adoption cost was about 3 days of initial tool development plus about 1 month for the practice to settle in. The operational conclusion: against six months of accumulated benefit, the adoption cost was small. That said, the ratios above are a single case from one team on one project; there is no guarantee they transfer to another team as is.


3.2.7 The AI–Schema Synergy, and Where It Ends

Once a schema is in place, the reliability of AI data generation jumps. The reason: the schema closes off, in advance, the ambiguous input ranges that give hallucination its opening. Given "make me 20 skills" with no schema, the AI invents plausible columns and fills in values incompatible with your sheets. With a schema, the same request comes back as rows that respect the seven defined columns, each constraint, and the FKs. Even when a violation slips through, as with row 1003 in 3.2.2, you point at one line, ask again, and you are done.

But the boundary is sharp. Balance values are never delegated to the AI. If the AI picks damage_coef "reasonably," it collides with the game's intent. The AI's share ends at laying out format-correct candidates quickly; "is 2.4 the right coefficient for this skill" is answered by a human. That does not mean AI is useless for balance — curve smoothness, outliers, and range statistics are things AI catches fast. Let the tools measure the numbers; let humans judge whether the numbers are right.


3.2.8 Common Mistakes and How to Avoid Them

Mistake Avoidance
Adopting a schema after 1000 rows have piled up New sheets always start with the $스키마
$스키마 and csv drift out of sync Bind the two to one source with schema-doc automation
Not stating FKs Without explicit FKs, the relation map and consistency checks are meaningless
Using proto Enum 0 for a semantic value 0 is _UNSPECIFIED (omission detection); semantic values start at 1
Schema docs read only by humans Markdown tables + uniform metadata so AI reads them too

Try It Yourself

setup 1. Pick the single most central sheet in your area (skills, items, or monsters). 2. Add a sheet named $스키마 to that Excel file, and write one line per column with the five elements (name, type, range, FK, description). Do this step yourself, by hand.

prompt (use AI only for the proto/csv drafts)

Convert the $스키마 below into a proto3 message and Enum. The first enum value must be _UNSPECIFIED = 0. Comment the reference target for FKs. Range constraints as comments only. Do not add columns that are not in the spec. (paste your $스키마 here)

Then:

5 csv rows satisfying the proto and $스키마 above. Do not produce rows that violate the constraints. If is_passive=true, damage_coef=0.

verify 1. Check the 5 rows the AI gave you against the schema, line by line. If a row violates a rule, ask again with "row N violates the rules; fix only that row" (rejection and retry are a normal part of the process). 2. Export the $스키마 to .md with schema-doc (or an equally simple Python script of your own) and confirm that the Excel definition and the document match. 3. If there are FKs, check once that the referenced PKs actually exist.


Solo Scale-Down

If you are starting alone, with no tools and no team, one Excel file and one text editor are enough.

  1. Create a $스키마 as the first tab of your sheet and write the column rules in five elements (15 minutes).
  2. Copy that spec as is and ask the AI for "a proto Enum + 5 csv rows" (10 minutes).
  3. Check the returned csv against the schema by eye, and fix the one violating row with a retry (10 minutes).
  4. Save the $스키마 text in a text editor as skill_schema.md. This is your first single source of truth.

When you move on to the next sheet, repeat the same four steps. Once 5–10 core sheets line up in the same order within a quarter, that is when automation like schema-doc becomes worth attaching.


Key Takeaways

Next Chapter Preview

3.3 Relation Map Visualization — Seeing Dependencies with Your Own Eyes

A new designer came to my desk during his first week at the company. "I'm about to touch the quest reward table — if I change this, what breaks?" I started to answer, pointing at my monitor, then stopped. The picture was in my head: RewardTable hooks into ItemTable, ItemTable hooks into ItemEffectTable, and above them QuestTable references the rewards... But the moment I put that picture into words, it lost its shape inside the listener's head. I drew seven boxes on a whiteboard. The arrows started to tangle. Thirty minutes later he nodded and went back to his seat — and the next day he came back with the exact same question.

That scene is what made me write this chapter. A systems designer carries a dependency graph in their head. The problem is that it exists only there. When the person changes, the picture disappears with them. I needed a tool to externalize the picture, and that is why I built gen_relation_map.py.

With 5\~10 data sheets, your head is enough. Past 30, human working memory cannot keep up. A project's sheet folder usually crosses that line early. A table that spells out in text which sheet depends on which never turns into a picture, no matter how carefully you read it. This chapter follows, start to finish, the worked process of auto-generating an interactive HTML relation map from foreign key (FK) relationships.


3.3.1 Four Problems a Relation Map Solves

Before building the tool, let me pin down what actually gets stuck when there is no relation map. Four scenes kept repeating.

Onboarding a new designer. A new designer books a meeting to learn the system structure. That is the scene above. Dependencies conveyed by mouth do not survive more than a few days in the listener's head. Click through a single relation map together, and more than half of the picture forms in the first meeting. The decisive difference from a hand-drawn whiteboard sketch is that the picture does not get erased — it stays in place.

Debating the impact scope of a change. A system change request comes in. "What does this affect?" A meeting gets scheduled, and even after a long discussion, one or two missed areas surface later. With a relation map, you click the node being changed and follow the inbound edges; the impact scope is right there in front of your eyes. The discussion only needs to settle whether each impact is real, and in what priority order.

Detecting dependency inversions. An L3 data sheet referencing an L1 system document is normal. The opposite direction — an upper Layer directly referencing a lower data sheet — is almost always a design flaw. In an FK list written out as text, humans cannot catch this inversion. In the picture, it shows up instantly as a single arrow whose Layer colors run the wrong way.

Finding orphaned sheets. Every so often you discover a sheet that nothing references. It is a leftover from an old design, or something that was retired on paper while the file stayed behind. It is like an unlabeled box rolling around in a corner of the office. You need the picture to spot that island.

What the four problems share: every one of them is only solved by seeing the structure with your eyes. Text and tables hit a wall.


3.3.2 Worked Transcript: From Data Sheets to a Relation Map

Now we follow the real thing. The input is one folder of data sheets; the output is a single interactive HTML page you open in a browser. In between, I record everything the AI did and every point where a human verified or rejected — nothing left out.

3.3.2.1 The Overall Flow

flowchart TD
    A[Data sheet folder
many xlsm/xlsx files] --> B[1. Scan: collect sheet and column headers] B --> C[2. Extract FK candidates
*_id / *Id / spec-sheet FK marks] C --> D[3. Match reference targets
column name → target sheet] D --> E{Human verification} E -->|reject false positives| C E -->|pass| F[4. Build graph
nodes=sheets, edges=FKs] F --> G[5. Assign Layer metadata
from schema-doc output] G --> H[6. Render HTML with pyvis] H --> I[relation_map.html
interactive in browser] I --> J{Human diagnosis} J -->|inversion/island/cycle found| K[Request design fix] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; class A,I data class B,C,D,F,G,H code class E,J,K human

The core is the human verification loop between steps 3 and 5. The machine lays down a draft of FK candidates, and a human weeds out the false positives. Skip this loop and the relation map becomes a plausible-looking but wrong picture.

3.3.2.2 Where the FKs Come From — Input Order

This tool's accuracy is decided by where it pulls its input from. The schema-first principle set in 3.2 applies as-is. The order of authority for FK information goes like this.

  1. The $스키마 sheet — the first source of truth for each data sheet. Type, Enum, and FK targets are specified per column. If an FK is written here, that is first priority.
  2. *.proto / Enum definitions — the schema exported via VBA (Excel's macro language). They fill in types when the spec sheet is empty.
  3. The actual csv output — the real data the sheet exports. Relationships missing from the spec sheet still surface as patterns in the data (for example, if every value in an npc_id column falls inside NPCTable's key range, it is effectively an FK).

Let me make one principle explicit here. The source of truth is the actual JSON/csv output, not the schema document. Even if the spec sheet says reward_id is an FK, if that column is empty in the real data or points at nonsense values, then the spec sheet is wrong. When the two disagree, the tool trusts the data side and records the disagreement itself in a report. This is why schema-doc is not treated as the source of truth.

3.3.2.3 Step 1 — Folder Scan and FK Candidate Extraction

The tool's first move is to open every xlsm/xlsx in the folder and collect sheet and column headers. This step is mechanical and not interesting. What is interesting is the rule for picking FK candidates. At first it was simple: if a column name ends in *_id or *Id, treat it as an FK candidate. But that single rule produced both false positives and misses.

This is where I brought in the AI. I handed over the entire list of column headers and had it infer the FK target sheets. Below are the full prompt I actually sent and the raw output.

Prompt: The following is a list of sheet names and column headers from game data sheets. For each column, infer whether it is a foreign key (FK), and if so, which sheet and which column it references. The usual naming convention is <singular form of target sheet>_id. However, if you are not confident, mark it "uncertain" and state why. Do not fill gaps with guesses.

QuestTable: quest_id, title, npc_id, reward_id, prev_quest_id, region_id
NPCTable: npc_id, name, region_id, faction_id
RewardTable: reward_id, item_id, count, currency_id
ItemTable: item_id, name, item_effect_id, grade
ItemEffectTable: item_effect_id, effect_type, value
RegionTable: region_id, name, parent_region_id

Claude's raw output (PK rows omitted; FK judgments excerpted):

QuestTable - npc_id → NPCTable.npc_id (presumably the NPC who gives the quest) - reward_id → RewardTable.reward_id - prev_quest_id → QuestTable.quest_id (preceding quest, self-reference) - region_id → RegionTable.region_id

NPCTable - region_id → RegionTable.region_id - faction_idUncertain: FactionTable is not in the input list. By the naming convention it should reference FactionTable.faction_id, but since that sheet was not provided, the target cannot be confirmed.

RewardTable - item_id → ItemTable.item_id - currency_idUncertain: CurrencyTable is not in the list. Currency types may also be managed as an Enum, in which case this may not be an FK. The spec sheet needs to be checked.

ItemTable - item_effect_id → ItemEffectTable.item_effect_id - grade → Not an FK (presumed Enum: grade values)

ItemEffectTable - effect_type → Not an FK (presumed Enum)

RegionTable - parent_region_id → RegionTable.region_id (self-reference, region hierarchy)

Note: the inferences above are based on column naming conventions; actual data values were not checked. In particular, columns that do not end in _id, such as count and value, were excluded from FK inspection — if hidden references exist there, they will be missed.

The human did three things with this output.

The lesson here is clear. Where the AI was most useful was not fast inference but the restraint to leave unknown slots marked "uncertain." Had it force-filled the blanks, faction_id would have been wired to some wrong sheet, and that false positive would have lived on in the relation map as a fake arrow, misleading the next new designer.

3.3.2.4 Step 2 — Graph Construction and Layer Assignment

Once the verified FK list is in hand, gen_relation_map.py builds the graph. Sheets are nodes; FKs are directed edges. Node size comes from counting inbound edges — how much do other sheets reference me. The more a sheet is referenced, the bigger the node: that is a hub of the system.

The Layer metadata is pulled from the Markdown schema documents generated by the schema-doc skill. The Layer coordinates defined in 3.1 (L0\~L4) are attached to each sheet as labels, and the tool reads them to color the nodes. This link matters. A relation map that does not know about Layers is just boxes and arrows; only when it knows the Layers can it diagnose "inversions" by color.

Here is the tool's internal structure as a code skeleton (core flow only).

# gen_relation_map.py (core flow excerpt)
from pyvis.network import Network

LAYER_COLORS = {          # Layer palette — standardized as a single atom
    "L0": "#2c3e50",      # meta/shared
    "L1": "#2980b9",      # system
    "L2": "#27ae60",      # content
    "L3": "#f39c12",      # data instances
    "L4": "#c0392b",      # derived/cache
}

def build_graph(fk_list, layer_map):
    net = Network(directed=True, height="900px")
    inbound = count_inbound(fk_list)          # aggregate inbound edges
    for sheet in all_sheets(fk_list):
        layer = layer_map.get(sheet, "L0")
        size = 10 + inbound[sheet] * 3        # bigger node for hubs
        net.add_node(sheet, color=LAYER_COLORS[layer],
                     size=size, title=sheet_tooltip(sheet))
    for src, dst, col in fk_list:
        # Detect Layer inversion: warning color when an upper Layer references a lower one
        edge_color = "#e74c3c" if is_reverse(src, dst, layer_map) else "#888"
        net.add_edge(src, dst, title=col, color=edge_color)
    return net

is_reverse is the small heart of this tool. If an edge's source sheet sits on a higher Layer than its destination (e.g., L1 → L3), the edge is treated as an inversion and painted red. When a person opens the picture and sees a red arrow, that is almost always a place that needs fixing.

3.3.2.5 Step 3 — HTML Rendering and the Resulting Structure

The last step is pyvis emitting the interactive HTML. Clicking a node pops a tooltip with that sheet's columns, Layer, and inbound count, and a search box filters by sheet name. This is exactly why it has to be HTML rather than a static PNG — once the node count passes a few dozen, the arrows in a static image tangle into something unreadable. You need to drag nodes apart with the mouse and narrow down to the area you care about by clicking before the pattern emerges.

Here is the structure of the relation map built from the example data above, redrawn as an SVG. Colors mean Layers; a red arrow (absent in this example) would mark an inversion.

RegionTable (L1) QuestTable (L2) NPCTable (L2) RewardTable (L3) ItemTable (L3) ItemEffectTable (L3) parent_region_id (self-reference)

Looking at node sizes, RegionTable is the most referenced (both Quest and NPC point at it). That is the hub. ItemEffectTable is a leaf node, so it stays small. For a new designer asking "where do I start if I want to understand this system," the answer is already in the picture, in node-size order.


3.3.3 What the Picture Diagnoses — Combined with Layers

We defined the Layer coordinates in 3.1. When this chapter's relation map lifts those coordinates into the visual realm, four diagnoses that were impossible with text or tables become possible on a single screen.

That said, this does not mean the picture catches every problem. The picture catches structural flaws. Whether a given FK is the semantically correct relationship — for example, whether npc_id really means "the NPC who gives the quest" or "an NPC who appears in the quest" — is not something the picture resolves. That is the human's domain judgment to make. The tool only sets the stage on which that judgment can operate.


3.3.4 Without Automatic Updates, It Rots

A relation map is not a make-once artifact. Sheets are added and changed every week. A relation map left to manual updates drifts from the real structure within a month or two, and a map that misleads is worse than no map at all. A team member burned once by a wrong picture stops looking at pictures — that is the most expensive failure.

So updates are wired to automatic triggers.

The generated HTML is auto-deployed to internal static hosting (the design portal). With nothing but a browser — no tool installs — everyone sees the same map. It is like a map kept permanently unfolded next to your desk. Whoever asks, you answer by pointing at the same picture together.


3.3.5 Common Mistakes and How to Avoid Them

Mistake Why it happens How to avoid it
Over 100 nodes, the picture tangles Cramming every domain into one screen Per-domain filtering, split views per group
Layer colors differ from tool to tool Palette redefined in every codebase Standardize the palette as a single atom (LAYER_COLORS)
FK detection only catches *_id, causing misses and false positives Relying on a single regex line Combine explicit FKs in the spec sheet with real-data value verification
The picture exists but nobody looks at it Not wired into the workflow Make attaching the picture to change requests and meetings mandatory
Built once, never updated, rots Relying on manual updates Automatic triggers are a must — manual goes useless within a month

In operating gen_relation_map.py, the row that burned me most often is the third one. Trust the *_id rule alone and you miss hidden references like count or value (the AI in 3.3.2.3 warned about this limit on its own), while falsely flagging the Enum grade as an FK. The verification loop that checks both the spec sheet and the real data is the answer to that row.


3.3.6 Try the Solo Scale-Down First

Trying to handle the company's entire set of data sheets at once is heavy, and you burn out before you ever get to show the value. Start small, with one folder from your own domain.

Try It Yourself

setup. 1. Pick one folder containing 5\~10 data sheets you own. 2. Install the dependencies with pip install pyvis openpyxl (for reading Excel, use the excel-reader skill or openpyxl). 3. First check whether FKs are specified in each sheet's $스키마 sheet. If not, just collect the column headers.

prompt. Gather the column header list and send the prompt from 3.3.2.3 as-is. The key is the last line — "If you are not confident, mark it uncertain, and do not fill gaps with guesses." That sentence is what blocks fake arrows.

verify. 1. Go through the AI's FK candidates line by line. Confirm the lines marked "uncertain" against the spec sheet or the real data. 2. Remove suspected Enum columns from the FKs (ones like grade and effect_type that look like FKs but carry no _id). 3. Check that the self-references (prev_*_id, parent_*_id) were caught correctly. 4. Draw the graph from the verified list, open it in a browser, and look for red arrows (inversions) and islands (isolation) with your own eyes.

Solo Scale-Down

If you have no time to build the tool, it is fine to start the first week with a single hand-written mermaid diagram. Write the FKs of 5 sheets directly into mermaid, in the format from 3.3.2.3. Take that one page into a meeting and show it — "this is our system's dependency graph" — and the value proves itself on the spot. Once the value is visible, the automation tool follows naturally. You can let go of the pressure to have a working tool from day one.

Scaling up flows naturally in this order — week 1: a hand-drawn mermaid of your own sheets → week 2: add Layer colors and clicking → month 1: automatic updates (git hook or nightly batch) → month 3: deploy to the internal portal → month 6: a unified relation map of all sheets.


3.3.7 Connecting to the Next Chapter

3.2 covered the inside of the sheets (schemas); 3.3 covered the outside (relationships). 3.4 layers AI-assisted prompt patterns on top. With schemas and relationships in place, it moves on to practical patterns for how AI assists with consistency checks and impact-scope extraction.


Key Takeaways

Next Chapter Preview

3.4 Prompt Patterns for AI-Assisted Systems Design

It was the week right before the alpha build. I added one new class row to the skill sheet and saved. What I did not learn until the build broke the next morning was that the buff ID this class referenced was a row someone had deleted the day before. Tracing back through the broken build to find the cause took two hours. Two hours I would never have spent if I could simply have asked, before that row was deleted, "Is this safe to delete?"

This chapter is about putting that question in the AI's job description. The core is not a knack for writing good prompts. It is making sure you never rewrite the same question from zero — pinning it down. Section 3.2 laid the schema, 3.3 the relation map. Both were the skeleton of the data. This chapter takes the questions a human throws at AI on top of that skeleton and hardens those questions themselves into assets.

Let me nail one thing down first. What AI produces is not an answer — it is a candidate. In every pattern in this chapter, the hand that makes the final decision stays on the human side to the very end.


3.4.1 Two Places Where Ad-Hoc Prompts Leak

When you first start using AI, you type a fresh natural-language prompt on the spot every time. Something like this.

스킬 시트 한번 봐줘. 외래 키 깨진 거 없나 확인하고,
이상한 거 있으면 알려줘. 아 그리고 쿨다운 음수인 것도.

(In English: "Take a look at the skill sheet. Check whether any foreign keys are broken, and tell me if anything looks off. Oh, and negative cooldowns too.")

This prompt leaks in two places.

First, the check items change every time. Today I remembered "negative cooldowns," but tomorrow I forget. The "duplicate PK (primary key)" check I ran yesterday is missing from today's prompt. A check that depends on human memory drops items exactly as often as the human's condition dips.

Second, the output format changes every time. Write the same intent as "check it," "inspect it," or "look it over," and the AI answers with a table one day and running prose the next. When the format wobbles, you cannot feed the results into any further automated processing.

The fix is to take the prompt out of your hands and put it in a drawer. The memo you used to scribble by hand each time becomes a labeled card you pull from the same drawer. That card is what this book calls a slash command (skill) and an atom.


3.4.2 Three Forms of Codification and How to Choose

There are three containers for codifying. What goes into which is decided by call frequency and definition stability.

Call frequency High ↑ Low ↓ Definition stability → Slash command (skill) Frequent · stable → /check-sheet One-word call, fixed output format Atom auto-injection (JIT) Frequent · core constraint → keyword trigger No memorizing — joins natural language Template file (.md) Occasional · large → call by file Easy to eyeball and edit Occasional · unstable definition → Do not codify yet (stay ad-hoc)

Work that is frequent and whose definition has hardened goes into a slash command. Work that is frequent but is really a constraint you must never forget goes into atom JIT auto-injection. Work that is occasional but large goes into a template file. And work whose definition is still shifting stays ad-hoc — do not codify it yet. You do not need all three from day one. Start with one or two slash commands and grow once the value shows.


3.4.3 Pattern ① Integrity Check — A Worked Transcript

Rather than explain in the abstract, I will walk one pattern through from start to finish: the pattern that automatically asks "Is this safe to delete?" before a single empty row gets deleted. Its name is /check-sheet. Inside it, the check items and the output format are codified.

The assets it stands on appear in the measured work logs embedded throughout this book. Data entry follows the schema-first principle (atom data_entry_schema_first). The entry order is the $스키마 (schema) sheet → Enum/*.proto (exported via VBA, Excel's macro language) → csv. And the source of truth is not the schema document but the actual JSON output (atom json_over_schema_doc_as_source_of_truth). The integrity check simply carries these two principles over into check rules.

setup — Inside the Codified Command

Open /check-sheet and this is the prompt body inside. This is the part you never have to type by hand again.

역할: 너는 게임 데이터 시트의 정합성 검사기다.

검사할 시트: {{sheet_name}}
참조 가능한 스키마: $스키마 시트 (컬럼별 타입·범위·FK 대상)
참조 가능한 정본: 같은 시트의 export JSON (스키마 문서보다 우선)

검사 항목 (이 순서로):
1. 외래 키 유효성 — FK 컬럼 값이 참조 대상 시트에 실재하는가
2. 범위 위반 — $스키마에 정의된 min/max를 벗어난 값
3. 필수 컬럼 누락 — required=true 컬럼이 빈 행
4. 중복 PK — 같은 PK가 둘 이상

출력 형식 (반드시 이 표 그대로):
| # | 행 | 컬럼 | 위반 종류 | 현재 값 | 수정 후보 | 확신도 |

규칙:
- 수정 후보는 제안만 한다. 절대 시트를 직접 고치지 않는다.
- 확신도는 높음/중간/낮음 셋 중 하나. 추측이면 낮음.
- 위반이 없으면 "위반 0건"만 출력하고 끝낸다.

(In English: the body assigns the role of an integrity checker for game data sheets, takes {{sheet_name}} as the target, points to the $스키마 sheet for per-column types, ranges, and FK targets, and to the exported JSON of the same sheet as the source of truth that outranks the schema document. It runs four checks, in this order: foreign key validity, range violations, missing required columns, duplicate PKs. The output is pinned to one fixed table — #, row, column, violation type, current value, suggested fix, confidence — and three rules close it out: suggest fixes only, never edit the sheet directly; rate confidence as one of high/medium/low, and if it is a guess, mark it low; if there are no violations, output only "0 violations" and stop.)

What matters most in this body is the final three rule lines: "do not edit directly," "rate your confidence," "confess guesses as low." Those three lines tie the AI to the seat that pushes candidates, not the seat that hands down answers.

prompt — The One Line the Human Actually Typed

/check-sheet skill_table

That is all. You pass one sheet name. You do not rewrite the four check items or the output table format. That is what codifying is worth.

Claude's Raw Output

Below is the output that came back when I actually ran it. I am not summarizing — the ambiguous parts and the spots where the AI hesitated are included as-is.

skill_table 정합성 검사 결과입니다. 총 4건의 위반을 발견했습니다.

| # | 행  | 컬럼        | 위반 종류    | 현재 값 | 수정 후보              | 확신도 |
|---|-----|------------|------------|--------|----------------------|------|
| 1 | 142 | buff_id    | 외래 키 위반 | 9007   | 9070 (오타 추정, 1글자) | 중간   |
| 2 | 142 | buff_id    | 외래 키 위반 | 9007   | 9012 (이름 "맹독" 유사) | 낮음   |
| 3 | 88  | cooldown   | 범위 위반    | -1     | 0 (음수 불가)           | 높음   |
| 4 | 215 | combo_next | 필수 누락    | (빈칸)  | 판단 보류               | 낮음   |

주의 사항 몇 가지를 덧붙입니다.

- 1번과 2번은 같은 행(142)의 같은 컬럼에 대한 서로 다른 후보입니다.
  9007이라는 ID는 buff 시트에 존재하지 않습니다. 9070은 숫자 한 자리
  차이라 오타일 가능성이 있고, 9012는 이름이 "맹독"으로 유사해서
  후보로 올렸습니다. 둘 다 확신은 낮습니다. 사람이 골라야 합니다.

- 4번(215행 combo_next 빈칸)은 애매합니다. $스키마에는 required=true로
  되어 있는데, export JSON을 보니 이 컬럼이 빈 행이 12개 더 있습니다.
  명세서가 최신이 아니거나, "콤보 종결 스킬은 combo_next를 비운다"는
  규칙이 있는데 명세서에 반영이 안 됐을 수 있습니다. 제가 단정할 수
  없어서 "판단 보류"로 두었습니다.

- cooldown 외에 범위 검사를 더 돌리고 싶으면 $스키마의 다른 컬럼
  min/max를 확인해야 하는데, 제가 받은 정보에 range_check 컬럼 일부가
  비어 있어서 cast_time 컬럼은 검사하지 못했습니다.

(In English: the check reports 4 violations. Findings 1 and 2 are two competing fix candidates for the same cell — row 142's buff_id is 9007, which does not exist in the buff sheet; 9070 differs by one digit and could be a typo (medium confidence), while 9012 is a buff with the similar name 맹독, "Deadly Poison" (low confidence) — a human has to choose. Finding 3: row 88's cooldown of -1 should be 0, negatives are not allowed (high confidence). Finding 4: row 215's blank combo_next is flagged but left as "judgment withheld" — the spec says required=true, yet the exported JSON shows 12 more rows with this column blank, so either the spec is stale or there is an unwritten rule that combo finisher skills leave combo_next empty. The AI also confesses a limit: part of the range_check information it received was blank, so it could not check the cast_time column.)

What makes this good output is not that it is tidy — it is that it is honest. It did not flatly declare finding 4 "required, so fill it in"; it exposed its own doubt — "the spec may not be current." It also confessed the limit that it could not see cast_time. If the AI had confidently said "fill it with 0," that could have been the wrong answer that breaks combo finisher skills.

verify — Human Verification and Rejection

Now the human rules on each line. "Verification comes bundled with the command" means this table comes out automatically — not that no human looks at it.

That last one is the heart of it. At the exact spot where trusting the AI's output as-is would have wrecked perfectly healthy data, the human rejection kicked in. At the same time, finding 4 exposed a defect in the check rule itself: the spec's required=true was wrong.

Re-Request — The Rejection Fixes the Rule

It does not end there. A rejection is input that makes the next check smarter. The human types on.

4번은 위반이 아니야. combo_finisher=true인 스킬은 combo_next가 비는 게
정상이거든. 이 예외 규칙을 검사 항목에 넣어서, 앞으로는 finisher 행의
combo_next 빈칸은 위반으로 잡지 마. 그리고 명세서의 required 정의가
틀렸다고 보이는 항목은 따로 "명세서 의심" 섹션으로 빼서 알려줘.

(In English: "Finding 4 is not a violation. Skills with combo_finisher=true are supposed to leave combo_next empty. Add this exception to the check items, so that blank combo_next on finisher rows is no longer flagged. And when a spec's required definition looks wrong, split those items out into a separate 'spec suspect' section.")

이해했습니다. /check-sheet의 검사 규칙에 예외를 반영하겠습니다.

추가된 규칙:
- combo_next 필수 검사 시, 같은 행의 combo_finisher 값을 먼저 확인.
  combo_finisher=true이면 combo_next 빈칸은 정상으로 처리(위반 제외).

신설 섹션:
- "명세서 의심" — export JSON의 실제 분포가 $스키마 정의와 어긋날 때,
  데이터가 아니라 명세서를 고쳐야 할 후보로 분리해 보고.
  (정본은 JSON이라는 원칙에 따라, 다수 데이터가 명세서를 거스르면
   명세서 쪽을 의심합니다.)

이 변경을 /check-sheet 명령 정의 atom에 기록해 둘까요?

(In English: Claude confirms the exception — when checking combo_next as required, first check the same row's combo_finisher; if true, a blank combo_next is treated as normal — and creates a new "spec suspect" section that reports cases where the actual distribution in the exported JSON contradicts the $스키마 definition, separated out as candidates for fixing the spec rather than the data, following the JSON-is-the-source-of-truth principle. It closes by asking: "Shall I record this change in the /check-sheet command definition atom?")

That final line completes the pattern. One rejection does not end as a one-off fix; it gets recorded in the command definition atom and propagates to the next person and the next check. This is the compound interest of codifying. You never filter the same mistake twice.


3.4.4 Pattern ② Change Impact Extraction — A Worked Transcript

The second pattern erases the two hours from the opening scene. Before changing one row or one decision, it first asks, "How far does this spread?" Its name is /impact. The relation map built in 3.3 (the FK-relationship HTML that gen_relation_map.py produces) is this pattern's fuel.

setup — The Codified Search Paths

Inside /impact, the answer to "where do you have to dig to find every impact without misses?" is codified as four paths.

역할: 너는 변경 영향 범위 추출기다.

변경 대상: {{target}}  (시트명·컬럼·atom_id·시스템 결정 중 하나)

탐색 경로 (넷 다 돌 것):
1. 관계도 인바운드 — gen_relation_map.py 출력에서 {{target}}을
   FK로 참조하는 시트들 (이 시트를 가리키는 화살표)
2. 본문 언급 — sources/manuscript 안에서 {{target}}의 ID·이름이
   직접 등장하는 문서
3. atom affects 관계 — {{target}}과 affects로 연결된 atom들
4. 역참조 — {{target}}을 [[wikilink]]로 가리키는 노트

출력 형식:
| 종류 | 이름 | 연결 근거 | 영향 강도 | 담당 |

규칙:
- 영향 강도는 직접/간접/약 셋이다. 직접 FK는 '직접', 본문 언급은 '간접', 추정은 '약'.
- 누락이 두려우면 약으로라도 올린다. 빠뜨리는 것보다 낫다.
- 우선순위나 작업 지시는 하지 않는다. 목록만 만든다.

(In English: the role is a change impact extractor; the target {{target}} is one of a sheet name, a column, an atom_id, or a system decision. It must walk all four search paths: 1) relation-map inbound — sheets that reference the target as an FK in the gen_relation_map.py output, i.e., arrows pointing at this sheet; 2) text mentions — documents under sources/manuscript where the target's ID or name appears directly; 3) atom affects links — atoms connected to the target via affects; 4) reverse references — notes pointing at the target via [[wikilink]]. The output is a fixed table — type, name, evidence of linkage, impact strength, owner. Rules: impact strength is one of direct/indirect/weak — a direct FK is "direct," a text mention is "indirect," an inference is "weak"; if you fear a miss, list it at least as weak — better than dropping it; no priorities, no work orders — produce the list only.)

prompt

/impact buff_id=9012

One line asks: "I am considering replacing the Deadly Poison buff (맹독, 9012) with a different effect — what gets shaken?"

Claude's Raw Output

buff_id=9012(맹독)의 변경 영향 범위입니다. 4개 경로를 모두 탐색했습니다.

| 종류    | 이름                | 연결 근거              | 영향 강도 | 담당      |
|--------|--------------------|----------------------|---------|----------|
| 시트    | skill_table         | buff_id FK 직접 참조    | 직접       | teammate_a |
| 시트    | monster_drop        | reward_buff FK 참조     | 직접       | teammate_b |
| 시트    | set_item_effect     | proc_buff FK 참조       | 직접       | teammate_a |
| 문서    | 독계열_밸런스_노트.md  | 본문에 "9012" 직접 언급  | 간접       | (저자)     |
| atom    | poison_stack_rule   | affects: 9012          | 간접       | (저자)     |
| 노트    | 상태이상_설계.md       | [[맹독]] 역참조          | 약       | (저자)     |

몇 가지 덧붙입니다.

- skill_table에서 9012를 참조하는 행은 3개입니다(암살자 클래스 스킬).
  교체하면 이 3개 스킬의 실제 효과가 바뀝니다. 가장 강한 영향입니다.

- monster_drop의 reward_buff는 제가 확신이 좀 약합니다. 컬럼 이름이
  reward_buff인데 일부 행은 buff가 아니라 item을 가리키는 것 같은
  값이 섞여 있어서요. 이 시트는 직접 한 번 열어 확인하시길 권합니다.

- "독계열_밸런스_노트.md"는 9012를 텍스트로 언급할 뿐 FK는 아니라
  간접으로 두었습니다. 다만 밸런스 의도가 적힌 문서라 실제 영향은
  강할 수 있습니다. 강도(간접)와 중요도가 다를 수 있다는 점 유의하세요.

- set_item_effect는 세트 아이템 4종에서 proc_buff로 9012를 씁니다.
  맹독을 빼면 이 세트들의 컨셉이 바뀝니다.

(In English: all four paths were searched and six entries came back — three sheets with direct FK references (skill_table via buff_id, monster_drop via reward_buff, set_item_effect via proc_buff), one document (독계열_밸런스_노트.md, a poison-line balance note that mentions "9012" in its text — indirect), one atom (poison_stack_rule, linked via affects — indirect), and one note (상태이상_설계.md, a status-effect design note linked only by a [[맹독]] reverse reference — weak). The AI adds caveats: 3 skill_table rows (assassin-class skills) reference 9012, the strongest impact, since replacing it changes those skills' actual effects; it is less confident about monster_drop because some reward_buff values seem to point at items rather than buffs, so it recommends opening that sheet yourself; the balance note is only a text mention, but since it records balance intent, its real importance may exceed its "indirect" strength; and 4 set items use 9012 as proc_buff, so removing Deadly Poison changes those sets' concepts.)

verify — Human Verification and Rejection

Here too, the AI's honesty helped the verification. Had it flatly declared monster_drop "direct," the impact would have been inflated and an unnecessary coordination meeting would have landed on the calendar. Because the AI handed the ball back to the human — "I recommend you check" — the strength got corrected on the spot.

Re-Request — From Extraction to Sharing

Once the list is verified, the next human move is delivering it to people.

직접 영향 받는 담당자별로 묶어서, teammate_a / teammate_b에게
보낼 짧은 공유 메모 초안 만들어줘. "9012 맹독을 X로 교체 검토 중,
당신 시트 N개 행이 영향" 형식으로. 결정은 아직 안 났으니 통보가
아니라 의견 요청 톤으로.

(In English: "Group the directly affected entries by owner and draft short share memos for teammate_a and teammate_b, in the form 'considering replacing 9012 Deadly Poison with X — N rows in your sheets are affected.' The decision has not been made, so set the tone as a request for opinions, not an announcement.")

담당자별 공유 메모 초안입니다. 통보가 아닌 의견 요청 톤으로 잡았습니다.

[teammate_a 앞]
맹독 버프(9012)를 다른 효과로 교체하는 안을 검토 중입니다. 확정은
아니고 의견을 먼저 듣고 싶습니다. 영향 범위:
- skill_table: 암살자 스킬 3개 행이 9012를 참조 (직접 영향)
- set_item_effect: 세트 아이템 4종이 proc_buff로 9012 사용
교체 시 이 스킬·세트들의 컨셉이 바뀝니다. 반대나 보완 의견 있으면
이번 주 안에 알려주세요.

[teammate_b 앞]
monster_drop의 reward_buff 관련해 확인 부탁입니다. 맹독 버프(9012)
교체를 검토 중인데, 영향 행이 1개로 보입니다(reward_buff 컬럼에
item ID가 섞여 있어 처음엔 더 많아 보였습니다). 실제로 이 1개 행만
맞는지 확인해 주실 수 있을까요?

두 메모 모두 "검토 중·의견 요청"으로 두었고, 확정 통보 표현은
넣지 않았습니다. 보내기 전에 교체 후보 효과(X 자리)를 채우셔야 합니다.

(In English: two memo drafts, both in opinion-request tone. To teammate_a: a replacement of the Deadly Poison buff (9012) is under review, nothing is confirmed and opinions come first; affected — 3 assassin-skill rows in skill_table referencing 9012 (direct impact) and 4 set items using 9012 as proc_buff in set_item_effect; the replacement would change those skills' and sets' concepts; please raise objections or suggestions within the week. To teammate_b: a verification request — the 9012 replacement is under review and only 1 monster_drop row appears affected (it looked like more at first because item IDs are mixed into the reward_buff column); could you confirm that this 1 row is indeed the only one? Claude notes that both memos stay at "under review, opinions requested" with no announcement phrasing, and that the replacement-candidate effect (the X slot) must be filled in before sending.)

The two hours from the opening scene disappear right here. You ask about the impact before the change, share it with the owners in advance, and coordination finishes before the next morning's build can break. And what the AI made remains, to the end, a memo draft — the finger that hits send is human.


3.4.5 The Skeleton Both Patterns Share

The two transcripts above follow the same skeleton. Every other pattern rides on it too.

flowchart LR
    A[Codified command
check items, output format, prohibitions] --> B[Human: one-line call
passes only the target] B --> C[Claude: raw output
candidates + confidence + self-doubt] C --> D{Human verification} D -->|accept| E[Apply] D -->|correct| F[Fix strength or values] D -->|reject| G[Rule defect found] G --> H[Update command definition atom] H -.applied to the next call.-> A classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; classDef fail fill:#fee2e2,stroke:#dc2626,color:#7f1d1d; class A,H data class B,D,F human class C ai class E pass class G fail

The key is the dotted line at the lower right. A rejection is not the end — it loops back as input that fixes the command itself. The rejected finding 4 in the integrity check became the finisher exception rule, and that rule was recorded in the atom and propagated to the next check. Without this feedback, you re-filter the same wrong answer every week.

Human hands remain in three places: the call (choosing the target), the verification (accept, correct, reject), and the rule improvement (feeding rejections back). In between, the AI does one thing only — pushing candidates.


3.4.6 The Remaining Patterns — Variations on the Same Skeleton

Carry the same skeleton over to other work and the patterns multiply. I will only mark their places here, without transcripts. They all follow the 3.4.5 skeleton as-is, so the key when building them is not to drop the "codified check items" and the "human verification slot."

Pattern One-line call Candidates the AI pushes Decisions the human keeps
GDD draft synthesis /gdd-new <system> Standard 9-section draft, [TBD] for the undecided Vision, priorities, deletions
State machine / Behavior Tree (BT) conversion /diagram-state Natural language → mermaid + reachability check State definitions, transition conditions
Interface conflict check /check-interface <GDD> Input/output and time-window conflict cases Priority rules
Balance calculation /balance-calc <sheet> <formula-atom> Curve values + diff against the current ones Formula, game intent
Retrospective work classification /retro-classify <period> Layer × discipline distribution + anomaly signals Classification fixes, interpretation

One thing to nail down about balance calculation. Even when a curve comes out numerically smooth, whether that smoothness matches the game's intent is a different question. You wanted the stretch right before the boss deliberately steep, and the AI shaves it flat as an "outlier" — that happens. So even with curve verification attached automatically, the last line of a balance calculation closes only after a human has held it up against the intent.


3.4.7 Five Operating Principles and a Convergence Point

As patterns multiply, you need operating discipline. The five principles below are not rules to memorize — they are design principles you build into the tools themselves.

Principle Why
One command = one job Smaller is easier to reuse and debug. Do not cram checking, fixing, and sharing into one /check
Bundle verification into the command Put confidence and evidence columns into the output table itself, to lighten the human verification load
Keep command definitions as atoms Like the rejection-to-rule loop in 3.4.3, record the why, the examples, and the change history in an atom
Measure usage frequency Commands called less than once a month are deprecation candidates. Cut by data
Human hands touch decisions only A command goes as far as generating candidates. Automated decisions are forbidden

One convergence point to close on. On one MMORPG project I ran, the slash commands that stayed stable in systems design converged over time to around 12. That is not a published standard — it is one project's observation (author's experience, unverified). But the direction is clear. You do not grow commands without limit; you add and remove one or two a month and stop at the number that fits in your head. A drawer with 100 labels is the same as a drawer with none.

Do not adopt everything at once. For the first month, codifying one weekly-repeated task into a slash command is enough. Once that one shows its worth, it spreads naturally to two, then three, the next month.


3.4.8 Closing Out Part 3

3.1 laid down the Layer coordinates of systems design, 3.2 the schema, 3.3 the relation map, and 3.4 put AI-assisted prompts on top of them. Here is how the week changes for a systems designer who has come through all four chapters.

flowchart LR
    M[Mon: GDD draft synthesis
4h → 1h] --> T[Tue: integrity check
half a day → 5 min] T --> W[Wed: onboarding with one relation map
5 meetings → 1 page] W --> Th[Thu: impact extraction
1h meeting → 10 min] Th --> F[Fri: retrospective auto-classification
manual → data-driven] F --> R[Reinvest the saved time
in design, review, player experience] classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; class M,T,W,Th,F ai class R pass

Grunt-work hours shrink, and those hours flow back into deep design and thinking about the player experience. Not refilling the saved time with more grunt work — that is the real reason for bringing in the tools.

Part 4, up next, is combat design. It is systems design's closest sibling, and the tools and patterns of 3.1–3.4 carry over as-is.


Key Takeaways


Try It Yourself

setup. Pick one check you repeat every week (for example, sheet integrity). Write down that task's 4 check items and an output table format, and codify them as a single slash command. Make sure the three rules ("do not edit directly / rate your confidence / confess guesses") go into the command body.

prompt. Call it by passing only the target, in one line.

/check-sheet skill_table

verify. Rule on the returned table row by row — accept, correct, or reject. When a rejection appears, that is not bad luck; it is a defect in the rule. Send one more line adding that exception to the command definition, so the next call never filters the same mistake twice.

Solo Scale-Down

If you have no team and no atom system, keep one text block in your notes app instead of a slash command. Title it "sheet check prompt." The contents: the 4 check items from setup above plus the 3 rules. Every time you run a check, copy the block, swap in the sheet name, and paste it to the AI. When something gets rejected, add the exception line to that note block yourself. Whether the tool is a slash command or a single note, the cycle — codify → call → verify and reject → update the rule — turns exactly the same.

Part 4 · Combat Design

4.1 The Combat Designer and the Layer Stack — Which Cell Does Game Feel Go In

Learning Goals for This Chapter (difficulty 🟡 practitioner · prerequisites: basic arithmetic and spreadsheet math): By the end, you will be able to decompose an abstract adjective like game feel into measurable signals, and assign coordinates for which Layer cell each of the combat designer's five deliverables sits in.

A build review meeting. The programmer pulls up the new skill they just hooked in. The character swings a sword; the enemy gets knocked back. Five people are watching. Someone says,

"Hmm... the hit feels a little weak somehow."

The person next to them nods. "Yeah, it's kind of flat."

The programmer asks, "What should I change, and by how much?"

Silence. None of the five people in the room can answer that question with a number. All five felt that "the hits feel weak," but nobody can say "take the hitstop from 3 frames to 5." The meeting spends 40 minutes trading adjectives — "make it heavier," "it lacks impact" — and ends with "let's look at it again in the next build."

This scene compresses every problem in combat design. It is the area players feel most directly, yet the moment you put that feeling into words, all you have left are adjectives. Adjectives can't be measured, and what can't be measured can't be tuned. The combat designer's first job is to pull those adjectives down into numbers.

This chapter decides which cell those numbers go in: where each of the combat designer's five deliverables sits on the Layer stack, and why those coordinates are the precondition for automation. The hands-on tools in 4.2, 4.3, and 4.4 move on top of these coordinates.

One Line for Non-Specialists. You don't need to memorize combat values or frame units in this part. The single thing to take away is this — "a request that travels as adjectives can be neither measured nor tuned." The moment you pull "make it heavier" down to "change what, to what value," collaboration starts to move — and that idea applies unchanged to vague feedback in any job outside games. Skim the five deliverables in 4.1.1 lightly; this one idea is all you need to carry forward.


4.1.1 The Five Things on a Combat Designer's Desk

Summed up in one line, the deliverables a combat designer owns cover "the entire process by which player input is converted into on-screen action." I split that into five chunks.

First, the combat Look & Feel spec. A document that translates abstractions like game feel, responsiveness, and weight into measurable values. This is the hardest deliverable in the discipline, and it becomes the yardstick by which the other four are judged.

Look & Feel breaks down into four signals.

Without this spec, the meeting-room scene repeats. With it, the tuning order comes out as "hitstop 3→5 frames, camera shake amplitude +20%."

Second, the skill, combo, and cancel system. The rules by which input is converted into action.

Third, character and monster AI. NPC behavior logic — Behavior Trees (BT), state machines (FSM (finite state machine)/HFSM), and decision tables. Monster behavior patterns, boss phase transitions, allied NPC cooperation, and crowd simulation all live here.

Fourth, the damage, resource, and cooldown formulas. The math that converts player choices into outcomes: damage coefficients, defense mitigation, criticals, elemental modifiers; resource (MP/energy/stamina) consumption and recovery curves; cooldown distribution.

Fifth, the animation control spec. The blueprint that determines how design intent actually looks in the build — animation graphs, BTs, IK hookups. This is usually a collaboration with programmers and animators, but if the designer doesn't supply a spec of intent, the intent breaks in the build. Hand over the materials without a blueprint, and a different house gets built.

The key point here is that all five meet on the same desk. Change the combo rules (second) and the DPS in the damage formulas (fourth) shifts, which in turn changes the perceived weight in Look & Feel (first). If it isn't explicit which deliverable is the input to which, a single change shakes five places at once. That's why we need coordinates.


4.1.2 Which Layer Cell the Five Deliverables Sit In

Now I place the five combat deliverables on the L0–L4 coordinates established in 2.3. This mapping is the spine of the chapter.

flowchart TD
    L0["L0 · Vision
'Action combat where the hits feel alive'"] L1["L1 · System Skeleton
Combo & cancel structure / Look&Feel spec / class skeleton"] L2["L2 · Content Flow
Per-chapter enemy pack curves / skill unlock order"] L3["L3 · Data Sheets
Damage coefficients, cooldown values, resource costs"] L4["L4 · Build Measurement
Measured DPS / actual combo routes / player feedback"] L0 -->|"received"| L1 L1 -->|"skeleton defines the flow"| L2 L1 -->|"spec anchors the values"| L3 L2 --> L3 L3 -->|"lands in the build"| L4 L4 -.->|"measurement → spec revision feedback"| L1 L4 -.->|"outliers → sheet adjustments"| L3 classDef vision fill:#2d3748,stroke:#1a202c,color:#fff classDef build fill:#c05621,stroke:#7b341e,color:#fff class L0 vision class L4 build

Restated as a table:

Layer Combat design deliverable Change frequency
L0 (received — vision: "action combat where the hits feel alive") Nearly fixed
L1 Combo & cancel structure / Look & Feel spec / class skeleton Slow
L2 Per-chapter enemy pack progression curves / skill unlock flow Medium
L3 Skill damage coefficient sheets, cooldown values, resource costs Fast
L4 Measured DPS in the build, actually viable combo routes, player feedback Every build

What makes combat design distinctive is that L4 carries more weight here than in any other discipline. In narrative design, the L1 spec is very nearly the final product; combat is different. Whether "the hits feel good" lives in the territory you can only know by picking up the controller in the build and watching the screen. Even if the spec says "hitstop: 5 frames," whether that actually feels weighty is confirmed only at L4. That's why simulation and automated measurement tools create their biggest value in this discipline (4.4).

But a large L4 doesn't make L1 any less important. Look at the dotted arrows. L4 measurements feed back into the L1 spec. Without a spec, the measurements lose their point of comparison. Only with a 5-frame spec can you get the diagnosis "measured 4 frames — 1 frame missing." The cycle of spec → build → measurement → spec revision passes through all five Layers. The combat designer keeps a hand on this entire cycle — the seat where more Layers spin simultaneously on one desk than anywhere else.


4.1.3 From Adjectives to Numbers — One Spine of the Real Work

That's enough abstraction. Let's go back to the meeting room and watch, as a worked transcript, the actual process of pulling "the hits feel weak" down into numbers. The tool is a single LLM. The output is not polished, and the parts the human rejected are left in as-is.

This is a reconstruction of the actual flow in which the combat task force (TF) of an MMORPG project I run as design director (hereafter "Project A") turned the Look & Feel of one new melee skill into a spec.

The Prompt (in Full)

Convert this skill's Look & Feel from adjectives into ms and frame values.
Five items: hit timing, hitstop, camera shake, effect synchronization, recovery.
Mark anything you're not sure of as "needs verification in this game."

Skill: greatsword overhead slash (heavy_overhead_slash)
Intent: "a heavy downward smash that staggers the enemy"
Frame rate: 60fps

(Translation: "Convert this skill's Look & Feel from adjectives into ms and frame values. Five items: hit timing, hitstop, camera shake, effect synchronization, recovery. Mark anything you're not sure of as 'needs verification in this game.' Skill: greatsword overhead slash (heavy_overhead_slash). Intent: 'a heavy downward smash that staggers the enemy.' Frame rate: 60fps.")

Claude's Raw Output (Not Summarized or Tidied)

The "weight" of the greatsword overhead slash is built mainly from the length of the hitstop and the recovery, plus the decay of the camera shake. Item by item:

1) Hit timing. The heavier the weapon, the longer the startup from input to activation needs to be for the weight to land. From input to the first visual response (the character beginning to raise the sword), 50–80ms (3–5 frames) is recommended. Note that the first visual response and the first hit response are different things. Time to impact depends on the motion length, so it can't be pinned down here — needs verification against this game's motion lengths.

2) Hitstop. The key variable for a weighty single blow. If light attacks sit at 1–2 frames, the greatsword overhead slash should be 4–6 frames (roughly 67–100ms). Past 6 frames, "weighty" tips over into "stuttery" and responsiveness suffers.

3) Camera shake. Amplitude around 1–2% of screen height, duration 0.2–0.3 seconds, and for decay, ease-out (fades quickly) rather than linear. For a heavy blow, a curve with a strong first oscillation that dies down fast feels natural.

4) Effect synchronization. VFX (hit effect), SFX (hit sound), and hitstop must all trigger on the exact same frame as the hit. Even a 1–2 frame mismatch reads as "flat." UI feedback (damage numbers) can lag 1–2 frames without harm.

5) Recovery. Weight comes at the price of long recovery. From the hit landing until the character returns to a vulnerable, actionable state, 20–30 frames (333–500ms) recommended. But this value varies widely with the game's overall combat tempo (fast action vs. deliberate soulslike) — needs verification against this game's baseline tempo.

In sum: "weight = long startup + 4–6 frames of hitstop + ease-out shake + same-frame VFX/SFX/stop + long recovery." These numbers are general recommended ranges, though — the actual sense of weight can only be finalized by testing it hands-on in the build.

Human Verification, Rejection, and Re-Request

The designer does not adopt this output wholesale. They stop at two points.

Rejection 1 — camera shake amplitude "1–2% of screen height." This game is mobile-first. On a small screen, 1–2% is barely visible, and there's the mobile motion-sickness issue on top of that. The designer rejects this recommendation and applies a house rule instead: "on mobile, express weight through emphasized hitstop, not shake." The LLM gave a general answer; it doesn't know this game's platform constraints.

Hold 2 — hitstop "4–6 frames." This is a hold, not a rejection. The range is right, but the exact value gets decided by feel in the build. The spec reads: "put 4 frames into the build as the default, create 5- and 6-frame variants, and compare all three by hand."

The re-request goes out like this.

This is a mobile-first project. Minimize camera shake and rewrite the spec
to express weight through hitstop, recovery, and SFX.
Lay out three hitstop variants — 4/5/6 frames — in a table for build comparison.

(Translation: "This is a mobile-first project. Minimize camera shake and rewrite the spec to express weight through hitstop, recovery, and SFX. Lay out three hitstop variants — 4/5/6 frames — in a table for build comparison.")

In this second output, the LLM produces a spec table that reflects the mobile constraints. That table goes into the build, and at the next build meeting the designer says, instead of an adjective, "the 4-frame variant is too light — adopt 5 frames." A 40-minute meeting becomes a 5-minute decision.

What This Transcript Shows

Three things. First, the LLM is good at producing a first draft that pulls adjectives down into numeric ranges — this is what breaks the silence in the meeting room. Second, the LLM does not know this game's constraints (mobile, tempo, motion lengths) — it can only give general recommendations, so rejecting and adjusting them is the human's job. Third, the LLM itself pinned down, twice, that "this can only be finalized by hands-on testing in the build" — even the tool knows that the final call on weight belongs to human hands at L4.


4.1.4 Four Places Where AI Recoups Its Adoption Cost

The transcript above showed only one place (spec writing). Across combat design as a whole, AI creates value in four places.

1) Simulation — the biggest value. Pre-compute DPS (damage per second) curves, combo routes, and resource consumption without a build. It is overwhelmingly faster than making a build and measuring by hand. We work through this directly with the simulate_dps simulator in 4.4.

2) Auto-generating state machines and BTs. Convert a natural-language description like "this boss enrages below 50% HP, and while enraged uses a 3-hit chain pattern" into a BT/FSM diagram. Accuracy is high — rule structures are territory LLMs handle well. It saves the time of moving the logic in your head onto a diagram.

3) Automated analysis of build captures. Automatically extract hit timing, combo success rates, and damage distribution from play footage. But this is the hardest of the four to implement (we weigh it honestly below).

4) Proposing balance adjustment candidates. Analyze each row of the data sheet, detect outliers and rough spots in the curves, and propose adjustment candidates. The human only chooses.

Of these four, automated capture analysis (3) has the widest gap between "can be done" and "can be done easily." Books often write "AI extracts everything from the footage automatically," but in practice it isn't that simple. Pixel-based computer vision on footage, off-the-shelf vision APIs, in-game telemetry logs — the accuracy and implementation-cost comparison of the three capture methods is covered authoritatively in 4.4; refer there. Here I'll state only the conclusion.

The most realistic path is in-game telemetry logs. Have the engine emit events directly, like "frame 1204: skill_overhead lands, damage 340, combo count 3." This is source data, so it's accurate, and it takes a single insertion of logging code. The LLM is used to read those logs and summarize them into natural-language reports ("resource efficiency holds up through 3-hit combos, then drops sharply from the 4th"). Footage remains a secondary aid — humans eyeball only the suspicious cases.

In other words, the realistic form of the "AI auto-analyzes the footage" vision is telemetry logs + LLM summarization, not pixel vision. That honest distinction is the starting point for the tool choices in 4.4.

And one thing that does not change across all four places: AI cannot make the final call that "the hits feel good." That belongs to the realm of player emotion, and the responsibility for that emotion stays with humans. AI only produces the supporting evidence for that emotional judgment, fast. Simulation numbers, BT diagrams, telemetry reports — all of it is material for a human to make the call by feel.


4.1.5 The Real Reason for Splitting the Coordinates — The Precondition for Automation

So far the reason given was the surface one: "split the deliverables across Layers and collaboration speaks a common language." I explained that combo rules go in L1 and damage sheets in L3 because their change frequencies differ. True, but not the whole story.

The essential reason for splitting the coordinates is that automation only works on top of them. The general thesis — Layer decomposition as the precondition for procedural generation and automation — was covered in 2.3, so here I narrow it to how that precondition plays out in three kinds of combat automation.

First, simulation only runs when "what is input and what can be changed" are kept separate. If the deterministic core (physics, hitboxes — the L1 skeleton) is mixed in with the changeable spec (damage values, cooldowns — the L3 sheets), the simulator cannot define its "space of change candidates." Core fixed, sheets variable — only with that separation can simulate_dps run a sweep like "raise the damage coefficient from 280 to 340 in steps of 20 and plot the DPS curve."

Second, automated capture analysis is only meaningful when action atoms are labeled. Only when the spec side has an atom labeled "this frame range is the hit phase of skill_overhead" can the signals extracted from telemetry logs be automatically cross-checked against the spec. Without labels, the log is a string of meaningless dots: "something landed at frame 1204."

Third, LLM combo-sequence generation only works when cancel rules and the input buffer are separated out into external documents. A bounded request like "within this character's 7 cancelable pairs and a 200ms input buffer, propose 10 five-hit combo sequences" is possible only when the cancel rules aren't frozen inside code but exist as standalone documents.

All three say the same single sentence. Mix the deterministic core with the spec and automation is blocked; separate them and automation opens up. The surface purpose of Layer decomposition is a common collaboration language; the essential purpose is to lay the preconditions for automated simulation, capture analysis, and LLM sequence exploration.

From Conservative to Progressive Application

Once this foundation is laid, combat operations evolve in two stages.

Conservative application — humans design, automation verifies. This is where most action and MMORPG combat operations are today. Humans write the combo/cancel specs directly; automation simulates DPS and resources, captures via telemetry, and produces "spec vs. measurement" comparison reports. Humans interpret the gaps, decide on spec revisions, and the cycle returns to spec writing. Design is human; simulation, capture, and comparison are automated.

Progressive application — AI proposes candidates, humans only adopt. The next stage. AI automatically enumerates 10–30 sequences within the cancel pairs and the input buffer; automation simulates each sequence's DPS and resources in parallel; the LLM attaches rankings and interpretation ("1st in resource efficiency, medium input difficulty"). What remains in human hands is one decision — "which of these sequences do we adopt as the signature?" — plus the director's calls on build integration and motion capture. Creating a sequence from zero and choosing among 30 are different orders of workload.

For progressive application to take root, three things must be in place: (1) deterministic simulation infrastructure that computes DPS, resources, and survival time in under a second without a build; (2) action atoms whose combo, cancel, and input-buffer rules are separated out into labeled external documents; (3) telemetry-based automated capture analysis. All three are direct products of the Layer decomposition described above.

Motion Capture Is an Irreversible Step — The Decision Gate

Finally, reversibility. The combat designer's review cycle mixes steps that can be undone with steps that can't, and knowing where that boundary lies matters.

Reversible ──────────▶ Decision gate ──────▶ Irreversible Combo & cancel spec revision Sim runs & reports (results freely discarded) Data sheet value tuning Build integration (dev) Partially reversible Decision gate Motion capture (signature actions) Capture studio, actors, reshoot costs Build integration (live) Hotfix costs, shifts in player perception

Motion capture is the thickest irreversible step in combat. Capture studio scheduling, actor booking, reshoot costs — all of it is expensive. So motion capture for signature actions proceeds only after simulation and automated capture analysis have run long enough to lock the sequence. Conservative or progressive, the decision gate sits just before motion capture and the live build. Every one of the combat designer's reviews must finish in the reversible steps to the left of that gate to be safe.


4.1.6 The Scene at the Studio — What Shrank

These are the changes Project A's combat TF measured over six months of running the coordinates and tools above. The figures below are rough averages pulled from the TF's operating records; the accurate way to read them is as the direction of perceived change, not as precise measurements.

Item Before adoption After adoption
Look & Feel meeting time 2 hours average (subjective debate) 30 minutes average (anchored to measurements)
Combo diagram authoring 1–2 hours per skill set 10 minutes per skill set
DPS curve verification Manual measurement after build (≈1 day) Simulation (≈10 minutes)
Balancing a new skill 3–4 build cycles 1–2 build cycles

The direction matters more than the numbers themselves. All four items moved from "subjective debate, manual measurement, build iteration" to "measurements, simulation, automated diagrams." Adjectives went down in the meeting room and numbers went up. That is the one sentence this whole chapter is trying to say — the combat designer's job is to build the bridge from the subjective (game feel, fun) to the objective (numbers, simulation), and AI is the tool that lays that bridge fast. The hand that decides "this feels weighty" at the end of the bridge is still human.


Key Takeaways


Try It Yourself — Pulling Adjectives Down into Numbers

setup. All you need is one LLM. Pick one skill you have on hand (new or existing). Write its intent as a single adjective line — "heavy," "nimble," "ponderous," something like that.

prompt. Fill your skill's details into the skeleton below.

You are a combat design assistant. Convert the Look & Feel of the skill
below into a "measurable numeric spec" — in ms, frames, and %, not adjectives.
Mark any item you're not sure of as "needs verification in this game."

Skill: [name]
Intent: "[one adjective line]"
Frame rate: [60fps, etc.]
Items: 1) hit timing 2) hitstop 3) camera shake 4) effect synchronization 5) recovery

(Translation: "You are a combat design assistant. Convert this skill's Look & Feel into a 'measurable numeric spec' — in ms, frames, and %, not adjectives. Mark any item you're not sure of as 'needs verification in this game.' Skill: [name] / Intent: '[one adjective line]' / Frame rate: [60fps, etc.] / Items: 1) hit timing 2) hitstop 3) camera shake 4) effect synchronization 5) recovery.")

verify. Ask two questions of every number in the output. (1) Does this value hold under this game's constraints (platform, tempo, motion lengths)? → If not, state the constraints and re-request. (2) Does this value need to be finalized hands-on in the build? → If so, write 2–3 variants into the spec instead of a single value and compare them in the build. Never adopt as-is any item the LLM itself flagged as "needs verification."

4.1.7 Solo Scale-Down

If you're building a game alone, you don't need all five deliverables and all five Layers. Do just two things at minimum. One, a one-page Look & Feel spec — for your 3–5 core actions, write down only hitstop, recovery, and synchronization as numbers. A memo written in adjectives will be unreadable even to you six months later. Two, pull combos and cancels out of the code into a single file — once the cancel pairs live as data, you can later ask an LLM to "propose 5 combos from these pairs." These two are the minimum coordinates that keep the door to automation open even in solo development.

4.2 Combat Look & Feel — Pinning Game Feel Down in Data

Five people are gathered in front of the meeting room monitor. The same build, the same skill, the same 30-second clip is playing on screen for the third time. The client programmer speaks first. "Looks fine to me." The artist crosses their arms. "It's weak. Something's missing." The designer next to them cuts in. "The effects are good, but it doesn't stick to your hands." The director watches for a long while and makes the call. "Hmm... let's go just a bit heavier."

And the meeting ends. Nobody wrote down how many milliseconds or how many frames "just a bit heavier" actually is. In the next build, the programmer implements their own understanding of "heavier," and the artist layers on their own. And the following week, in the same meeting room, watching the same clip, the same conversation repeats.

Hit feel, game feel, Look & Feel. These are the most frequently used and least defined words in combat design. Everyone believes they know what they mean, but everyone's mental definition differs, so the discussion ends with nothing to show for it. This chapter is about decomposing that "feel" into measurable numbers — about pulling game feel down from abstraction into data.


4.2.1 Combat = Look & Feel, Systems = Behavior

Let me draw a boundary first. Combat design splits into two broad branches.

This chapter covers only the latter. Whether the damage is 100 or 120 has no direct bearing on game feel. What matters is how the player perceives the moment that 100 damage "lands." Even with an identical damage formula, different hit timing and hitstop make it feel like a completely different game.

Let me be honest about one thing up front. Hit feel is not completed by three numbers alone. The acceleration and deceleration curves of the attack motion (animation), the reaction on the receiving end (hit reaction, hitstun), and the squash-and-stretch deformation that Japanese action games of the '80s and '90s loved — exaggerating the character at the moment of impact with stretching, squashing, afterimages, and smears — all have to come together before "I got hit" registers as a single, unified sensation. What this chapter focuses on pulling down into measurable numbers is three of those axes. Motion, reaction, and deformation are areas where animators and artists have the bigger hand, so they belong to later chapters and the art part; here the weight is on the three axes a game designer can pin down in a spec and verify in a build. Those three axes break down like this.

The Three Measurable Axes (not all of game feel) Hit timing Input → response When does it respond? Unit: ms "Is the response fast?" Hitstop Moment of impact How long time freezes Unit: frame "Does it feel heavy?" Effect sync VFX·SFX·UI· camera·rumble Unit: frame offset "Does it fire as one event?" + + These three are what gets measured — motion, hit reaction, deformation are art/animation territory (separate)

When someone in the meeting room says "it's weak," that weakness comes from one of the three. Is the response late (timing)? Is there no sense of impact (hitstop)? Are things firing out of step (synchronization)? Decompose the question into the three axes, and "it's weak" finally becomes a sentence you can fix.

That said, the cause of "weak" does not always live in these three axes. Here is the full list of what makes up Look & Feel, with a line drawn around what this chapter takes responsibility for.

Look & Feel component What it is In this chapter
Hit timing ms from input to first response. The first suspect Measured and specced (axis 1)
Hitstop How long time freezes at the moment of impact to give weight Measured and specced (axis 2)
Camera shake The screen recoil that accompanies a hit Measured and specced (part of axis 3)
VFX/SFX timing Whether effects and sound sync to the hit frame Measured and specced (part of axis 3)
Attack motion (animation) Acceleration/deceleration of the swing, the curves of anticipation and follow-through Mentioned (art/animation territory)
Hit reaction, hitstun The flinch and stagger on the receiving end Mentioned (next chapters, art part)
Deformation (squash and stretch) Exaggerated stretching and squashing at the moment of impact (afterimages, smears) Mentioned (art part)
Controller rumble Physical feedback delivered to the hands Measured and specced (part of axis 3)

The top four bundle into this chapter's three axes and become the targets of measurement and specification. The middle three (motion, reaction, deformation) must not be missing, but they are art/animation territory that a designer alone cannot close out with numbers, so I only make clear that they exist. If the motion is stiff or the target just stands there unfazed, hit feel will not land even with all three axes dialed in.


4.2.2 Hit Timing — From Input to Response

There is a reason timing comes first among the three axes. When players doubt the game feel, the first thing that snags is the sense that "the response is late," and no matter how flashy everything else is, sluggish input brings it all down in that instant. So timing gets fixed first.

The first axis of game feel is time. From the instant the button is pressed (0ms) to the instant the screen first responds — how many ms does it take? Humans are astonishingly sensitive to this latency. The difference between 60ms and 120ms is something "you can't explain in words, but your hands know."

A single attack is not one simple point; it is multiple events laid out across time. Put one basic attack on a timeline and it looks like this.

Basic attack, first hit — timeline (warrior / skill_id 1001) 0ms 100 150 250 350 Input (0) Casting motion 0~100 Hitbox 100~150 Visual effects 100~250 (fade after 50ms delay) Damage applied 110 (10ms after visuals → perceived as simultaneous) Recovery 150~350 (until next input is accepted)

The most important number in this diagram is the 100ms at which the hitbox first turns on. It means hit detection for the attack begins 100ms after the button press. This value determines the perceived speed of the game feel.

Recommended ranges vary by genre and character, but rough baselines exist.

Type Recommended input→response Notes
Instant response (light attacks) 60\~120ms The core range for attacks that "stick" to your hands
Heavy response (big skills) 200\~400ms Deliberate startup for the sake of weight
Charging (long charge-up) 500\~2000ms Intentional wait, handled separately

These ranges are not absolute standards. As an author's estimate (unverified): casual mobile tends to drift about ±50ms toward input leniency, while console fighting games tend to tighten further. What matters more than the numbers themselves is that the team shares a baseline — "we agreed our light attacks are 90ms." Only with a baseline can you look at a build and call it "right" or "wrong."

But there is a trap here. The human eye cannot tell 90ms from 110ms. At 60fps one frame is about 16.67ms, and this 20ms difference is barely more than a single frame. Whether "feels a bit slow?" in the meeting room is right or wrong can never be settled by eye. That is why measurement is needed.


4.2.3 Extracting Timing from the Build — An Honest Comparison

The spec says "hitbox 100ms." How do you confirm the build actually turns it on at 100ms? Automation splits three ways (video analysis, off-the-shelf vision tools, in-game telemetry); the precision and difficulty comparison of the three is treated canonically in 4.4. Here I will state only the conclusion. The first thing to lay down in practice is in-game telemetry. The reason is simple. Rather than inferring from video which frame the VFX "appeared" on screen, having the code print a single [HITLOG] line on the frame it fires the OnHit event is incomparably more accurate and cheaper. Save video analysis for external footage with no input overlay (e.g., analyzing a competitor's game); for our own build, telemetry goes in first.

A telemetry log looks like this.

[HITLOG] frame=6  t_ms=100  evt=hitbox_on    skill=1001 char=warrior
[HITLOG] frame=6  t_ms=100  evt=vfx_trigger  skill=1001
[HITLOG] frame=6  t_ms=100  evt=sfx_trigger  skill=1001
[HITLOG] frame=7  t_ms=117  evt=damage_apply skill=1001 dmg=124
[HITLOG] frame=7  t_ms=117  evt=ui_dmgnum    skill=1001
[HITLOG] frame=6  t_ms=100  evt=cam_shake    skill=1001 amp=0.4

The designer's job is to check this log against the spec, line by line. It is mostly unit conversion and mechanical matching — a human repeating it by eye gets tired and makes mistakes, but an LLM does not tire. In the next section, I put it to work.


4.2.4 Worked Transcript — Checking the Telemetry Log Against the Spec

I pasted in both the spec yaml and the build's telemetry log, and told the AI to compare the two and find where they diverge. Below is the full prompt, Claude's raw output, and the entire human verification process. Nothing is summarized; it is reproduced as is.

First Prompt

Below are the 'hit timing spec' for one combat skill and the telemetry
log pulled from the actual build. Organize the items where the build diverges from the spec into a table.

[Spec]
character: warrior
skill_id: 1001
fps: 60
events:
  hitbox_on:    {expect_ms: 100, tol_frame: 0}
  vfx_trigger:  {expect_ms: 100, tol_frame: 0}
  sfx_trigger:  {expect_ms: 100, tol_frame: 1}
  damage_apply: {expect_ms: 110, tol_frame: 1}
  ui_dmgnum:    {expect_ms: 100, tol_frame: 2}
  cam_shake:    {expect_ms: 100, tol_frame: 1}

[telemetry log]
[HITLOG] frame=6  t_ms=100  evt=hitbox_on    skill=1001
[HITLOG] frame=6  t_ms=100  evt=vfx_trigger  skill=1001
[HITLOG] frame=8  t_ms=133  evt=sfx_trigger  skill=1001
[HITLOG] frame=7  t_ms=117  evt=damage_apply skill=1001
[HITLOG] frame=9  t_ms=150  evt=ui_dmgnum    skill=1001
[HITLOG] frame=6  t_ms=100  evt=cam_shake    skill=1001

60fps (1 frame = 16.67ms). Convert expect_ms to frames, compare against
the log's frame values, and mark anything whose difference exceeds tol_frame as FAIL.

Claude's Raw Output (Verbatim)

First, I will convert the spec's expect_ms values to frames. At 60fps, frame = round(ms / 16.67).

Here is the comparison against the actual frames in the log.

Event Expected frame Actual frame Diff (frames) Tolerance Verdict
hitbox_on 6 6 0 0 PASS
vfx_trigger 6 6 0 0 PASS
sfx_trigger 6 8 +2 1 FAIL
damage_apply 7 7 0 1 PASS
ui_dmgnum 6 9 +3 2 FAIL
cam_shake 6 6 0 1 PASS

Two items exceeded their tolerance. - sfx_trigger: the sound is 2 frames (about 33ms) behind the visual hit. Exceeds the 1-frame tolerance. Players are likely to feel that "the sound is slightly out of sync." - ui_dmgnum: the damage number appears 3 frames (about 50ms) late. Exceeds the 2-frame tolerance. This can come across as "I landed the hit, but the number shows up late."

damage_apply at frame 7 matches the spec's frame 7 (converted from 110ms), so it is normal. However, the fact that damage_apply is 1 frame after hitbox_on (frame 6) may need confirmation as to whether this is intended design.

Human Verification — I Rejected One Cell

I did not take the output on faith. There were two verification points.

Verification 1 (rounding boundary — passed). The 110ms conversion for damage_apply looked suspicious. 110 / 16.67 = 6.598…. Rounded, 7 is correct. But what if the spec's intent had been "frame 6 — the same frame as 100ms"? I checked directly: the spec deliberately set damage_apply: expect_ms 110, a 10ms offset from vfx (100). So the frame 7 conversion was accurate. The AI passed.

Verification 2 (the last sentence — rejected). The AI's closing remark — "the fact that damage_apply is 1 frame after hitbox_on may need confirmation as to whether this is intended design" — was noise. The spec already states the intent with expect_ms 110, and the AI turned that into a "needs confirmation" question. Given the spec, the answer should be closed within the spec; instead it manufactured an unnecessary question. I did not adopt that sentence.

I re-requested, including the reason for the rejection.

Second Prompt (Re-request)

damage_apply has its intent pinned in the spec as 110, so drop the 'needs confirmation' comment.
// meaning: do not ask back about items whose answer is closed within the spec
Keep only the two FAILs (sfx, ui_dmgnum), and instead of asserting causes,
give me a checklist of 'code/resource locations to check'.

Claude's Second Output (Verbatim)

Here is the verification checklist for the two FAIL items. No causes are asserted; only the places to check are listed.

sfx_trigger (+2 frames late) - [ ] Is the sound trigger bound to the same event as hitbox_on, or is it hooked to a separate animation notify? - [ ] Does the audio asset itself have leading silence — an empty stretch at the head of the file can present as latency - [ ] Was the playback call issued on the same frame but pushed back by output latency (platform audio buffer)?

ui_dmgnum (+3 frames late) - [ ] Is the damage number UI structured to update on a separate tick after damage_apply (frame 7)? - [ ] Is UI updating batched at the end of the next frame (deferred render queue)? - [ ] Is there an extra delay value on the start frame of the number pop-up animation?

What can be asserted: both items exceed the spec's tolerance and are later than the visual hit (frame 6). What cannot be asserted: which of the causes above is responsible — unknowable until code/assets are checked.

This second output I adopted as is. Not inventing causes, and converting them into "places to check," was exactly the form I wanted. I handed the checklist directly to the sound owner and the UI owner. For sound, 33ms of leading silence in the audio asset was the culprit (checklist item 2). For UI, it was the next-frame update structure (item 1).

This is where the division of labor becomes clear. The AI mechanically compares spec against log and catches the FAILs; the human (a) rejects the unnecessary counter-questions the AI generates and (b) confirms the real cause of each FAIL in the code. Ask the AI to assert causes and it will produce plausible lies, so it is safer to stop it at "places to check."


4.2.5 Hitstop — The Weight of a Hit

The second axis is the stop. The effect of freezing — or briefly slowing — game time at the very instant a hit lands. This determines the intensity of the "I connected" sensation. It is the most powerful game-feel tool in fighting games and action RPGs. Too long and it feels sluggish; too short and there is no weight.

Recommended ranges (at 60fps) are below. These figures are rough conventions in action games; absolute values are tuned per game.

Type Recommended frames Converted
Light hit 1\~2 frames 16\~33ms
Medium hit 3\~5 frames 50\~83ms
Heavy hit (finisher) 6\~12 frames 100\~200ms
Critical / weak-point hit above values + 2\~3 frames

Each character and skill must get different values. Give everything the same number and the weight differences vanish — every attack eventually converges on the same tone. This is the point that leads into one of 4.2's three Key Takeaways.

"Who stops" is also a design choice.

Option Effect Fits
Attacker only Weight on the attacker's side; the victim's knockback/knockdown continues Action
Victim only Victim pauses briefly; attacker moves freely Combo-friendly
Both Strongest sense of weight Fighting-game tradition

The spec is entered like this.

character: warrior
skill_id: 1001
hit_stop:
  attacker: 2          # frames
  victim: 4
  critical_multiplier: 1.5   # 1.5x on critical (rounded)

Hitstop is the axis where game-feel verification is hardest — spec numbers alone cannot settle "right or wrong." Telemetry can catch whether it actually stopped for 4 frames, but whether 4 frames is appropriate has to be judged by a human, hands on the build. The AI guarantees conformance to the spec; the human judges whether the spec value itself is right.


4.2.6 Effect Synchronization — Does Everything Fire on the Same Frame?

The third axis is simultaneity. When VFX (visual effects), SFX (sound), UI (damage numbers), camera (shake), and rumble (controller) all start on the same frame, the player's brain binds them into "one event." Even a 1\~2 frame misalignment reads as "something's off"; at 3\~5 frames it reads as "looks like a bug." The sfx 2 frames late and the ui 3 frames late that FAILed in the earlier worked transcript were exactly this axis's problem.

The five synchronization targets and their tolerances:

Element Trigger point Tolerance
VFX (visual effects) Hit frame ±0 (must be simultaneous)
SFX (sound) Hit frame ±1 frame (16ms)
UI damage numbers Hit frame ±2 frames
Camera shake Hit frame ±1 frame
Controller rumble Hit frame ±2 frames

The key point is that all five belong in the spec. The common mistake is to spec only the VFX and leave the other four to "they'll line up somehow." If it is not in the spec, build verification has no baseline — even if telemetry captures it, there is nothing to compare against. Only when all five are in the spec does the automated comparison close.

For the automated comparison, the worked transcript from the previous section applies as is. If the telemetry log records the trigger frames for all five, the AI checks them against the spec and reports only the items beyond tolerance. No human has to eyeball 100 skills on every build. What remains separate, though: the AI catches deviations from the spec, while the human catches the territory where "the spec all passes, but it still doesn't feel right."


4.2.7 From Spec to the Next Build — The Full Loop

Now to connect the pieces into a single flow. Once this loop starts turning, the meeting room's "just a bit heavier" gets translated into "hitstop victim 4→6 frames."

flowchart TD
    A["Write the spec
(designer: yaml)"] --> B["Implement in the build
(programmers & artists)"] B --> C["Run the build + telemetry log
[HITLOG] auto output"] C --> D["AI auto-comparison
spec yaml vs telemetry"] D --> E{"Any FAIL beyond
tolerance?"} E -->|yes| F["FAIL items + places-to-check list
(AI, no cause assertions)"] F --> G["Owner confirms the real cause
in code/assets (human)"] G --> B E -->|no| H["Designer: 'feel' review
spec passes, but does it feel right? (human)"] H -->|needs tuning| A H -->|OK| I["Keep as input for the next milestone retrospective"] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; class C code; class D,F ai; class A,G,H human; class I pass;

In this loop, the boxes the AI owns (D, F) and the boxes humans own (A, G, H) divide cleanly. The AI is strong at mechanical comparison and checklist generation; humans are strong at setting baselines, confirming causes, and making the final call on feel. Automation does not remove the human — it releases them from the half-day of "counting frames by eye" they used to spend each cycle, so they can focus on the feel alone.

The last box in the loop (I) matters. This measurement data is not used once and thrown away; it feeds back in as input for the next milestone retrospective. When a pattern like "last quarter's game-feel FAILs clustered in sfx synchronization" survives as data, next quarter you fix the audio pipeline first.


4.2.8 Where Measurement Cut the Debate Short — Operational Observations

These are the changes I observed over roughly six months of running the loop above on a mobile MMORPG project where I served as director (hereafter "Project A"). The figures below are not precision instrumentation; they are the author's operational observations (estimates included), based on meeting-minute timestamps and build verification records. Read them for direction and ratio. Do not cite the absolute values as a benchmark.

Item Before After Nature
Time per Look & Feel meeting Dragged on Cut to less than half From meeting minutes; perceived
Build verification (many skills) Nearly a full day Sharply reduced Effect of automated telemetry comparison
Resolving "weak hit feel" feedback Several build cycles 1\~2 cycles Small sample; direction only
Spec-to-build conformance Around half Mostly conformant Became measurable after telemetry

The essence is the qualitative change, not the numbers. When "it's weak" came up in the meeting room, the immediate follow-up became "which axis? Timing? Hitstop? Sync?" — and when no answer came, we pulled up the telemetry together. That a measurable, objective standard gave the debate a destination is the biggest change of those six months. "Just a bit heavier" walked out of the meeting room undefined far less often.


4.2.9 Common Mistakes and How to Avoid Them

Mistake Avoidance
Verifying the build before there is a spec Spec first. Without a baseline, "right or wrong" is impossible
Reaching for video analysis first For our own build, telemetry first. Video analysis is for external footage only
Having the AI assert FAIL causes Stop at the "places to check" checklist. A human confirms the cause in code
Speccing only VFX and dropping the other four All five (VFX, SFX, UI, camera, rumble) go in the spec
Copying one character's spec to every character Differentiate per character and skill. Identical values converge game feel into one tone
Spamming hitstop on every hit Only on meaningful hits. Overused, it feels sluggish
Accepting the AI's counter-questions as is Reject "needs confirmation" comments on items already closed within the spec

Try It Yourself

This is the procedure for laying down a minimal version of this loop in your own project.

setup 1. Pick one skill to verify (a basic attack is recommended). 2. Plant one log line at each of the six trigger points in code (hitbox_on, vfx, sfx, damage_apply, ui_dmgnum, cam_shake): [HITLOG] frame=X t_ms=Y evt=... skill=.... 3. Write the spec yaml (an events block with expect_ms + tol_frame). Use this chapter's first prompt example as a template.

prompt 4. Run the build once and collect the telemetry log. 5. Paste the spec yaml and the telemetry log together into the AI and instruct it: "Compare the build against the spec using 60fps frame conversion and give me only the FAILs as a table. Do not assert causes — give a 'places to check' checklist." (the form of this chapter's second prompt)

verify 6. Recompute one cell of the AI's frame conversion yourself (ms / 16.67, rounded). If even one cell is wrong, distrust the whole thing. 7. If the AI asks back about items closed within the spec, or asserts causes, reject and re-request. 8. Hand the FAIL checklist to the owners and confirm the real causes in code and assets.

Solo Scale-Down

If you are a solo developer with no team and no telemetry infrastructure, scale it down like this. Screen-record the build at 60fps, with a key input overlay turned on so the moment of input is visible. Record one use of the skill you want to verify, then count frames yourself in a video editor: the frame the button was pressed, and the frame the screen first changed. Give the two frame numbers and the spec's expected value to the AI and say "convert to ms at 60fps and compare against the spec" — and even without telemetry, the one core axis (hit timing) gets verified. Full five-way synchronization is out of reach, but pinning down even the single input→response axis moves half of the game-feel debate onto objective ground.


The next chapter moves from a single hit to a sequence of hits. Combos, cancels, input queues — the rules that make one hit flow naturally into the next.


Key Takeaways

Next Chapter Preview

4.3 Combos, Cancels, and Input Queues — Enumerate the Paths and Verify Them

Combat designer B was standing at the meeting-room whiteboard, drawing boxes with a marker. Basic 1, Basic 2, Basic 3, and a heavy-attack branch peeling off to the side. Around the seventh arrow, someone asked: "So after the heavy attack into the launcher, if you cancel into dodge, can you get back to Basic 1?" B stopped the marker. The graph on the whiteboard didn't show that path. Whether it could be drawn and simply hadn't been, or whether the rules made it impossible — B couldn't answer on the spot either.

This is the real problem of combo design. In your head, a combo looks like a simple trunk: 1-2-3, then branch into the heavy attack. But once cancels and an input queue enter the picture, the trunk becomes a graph. Add just a few cancel edges to six nodes and the paths you can actually walk multiply into dozens. A human cannot unfold all those dozens of branches mentally. So the balance incident — "this path is too strong" — gets discovered only after it has gone into a build.

This chapter has one goal: a workflow that enumerates every combo path automatically, without drawing them by hand, and verifies each one. Turn rules written in natural language into a spec, enumerate the paths from the spec, and run the enumerated paths through simulation. Along the way I'll show, raw and unedited, how far the AI takes you and where it lies.


4.3.1 A Combo Is a Graph, Not a Table

Write a combo as a table and it reads like this: "Basic 2 follows Basic 1, Basic 3 follows Basic 2." Neat rows and columns. But this table lies, because a table assumes a straight line. In actual combat, the player branches from Basic 2 into the heavy attack, cancels the heavy attack into a dodge, and presses Basic 1 again right after the dodge. Those branches and cycles hide between the rows of the table.

So the true shape of a combo is a directed graph. Actions are nodes; connections are edges. Each edge carries an input window (when input is accepted) and an input key. Nodes carry a duration in frames, and some nodes carry a bonus condition (a damage multiplier that applies only if specific nodes were visited along the way).

Here is one basic combo set for a warrior character drawn as a graph — six nodes, cancel branches included.

Basic 1 (21f) Basic 2 (24f) Basic 3 (30f) Finisher ×1.5 Heavy (33f) Launcher (28f) Dodge (18f) 10~21f 12~24f 14~30f Heavy 6~24f Dodge cancel Re-enter Basic 1 after dodge

Two things make this decisively different from the whiteboard. First, every edge carries an explicit input-window frame range. The branch label "Heavy 6\~24f" means the heavy-attack input is accepted from frame 6 through frame 24 after Basic 2 starts. Second, there is the dashed dodge-to-Basic-1 re-entry edge — the very path B couldn't answer about in the meeting room. Make it explicit in the graph, and "exists / doesn't exist" becomes unambiguous.

Drawn by hand, this graph is six nodes and seven or eight edges. With twenty characters and three or four combo sets per character, you are looking at hundreds of graphs. Hands cannot keep up. So you write the graph down as a text spec, and generate both the picture and the verification from it automatically.


4.3.2 Humans Read the Spec, Machines Parse It

Let's move the graph above into a YAML spec. The core is three blocks: nodes (nodes), edges (edges), and bonuses (bonuses). Cancel rules are treated as just another kind of edge — interrupting an action to go to a different node is, in the end, an edge.

# warrior_basic_chain.yaml
character: warrior
combo_id: basic_chain

nodes:
  - { id: basic_1,  name: 기본1,   duration_frames: 21 }
  - { id: basic_2,  name: 기본2,   duration_frames: 24 }
  - { id: basic_3,  name: 기본3,   duration_frames: 30 }
  - { id: heavy,    name: 강공격,  duration_frames: 33 }
  - { id: launch,   name: 띄우기,  duration_frames: 28 }
  - { id: dodge,    name: 회피,    duration_frames: 18, cancels_recovery: true }

edges:
  - { from: basic_1, to: basic_2, input: light, window: [10, 21] }
  - { from: basic_2, to: basic_3, input: light, window: [12, 24] }
  - { from: basic_2, to: heavy,   input: heavy, window: [6, 24] }
  - { from: heavy,   to: launch,  input: heavy, window: [10, 33] }
  - { from: heavy,   to: dodge,   input: dodge, window: [0, 33], type: cancel }
  - { from: basic_3, to: dodge,   input: dodge, window: [0, 30], type: cancel }
  - { from: dodge,   to: basic_1, input: light, window: [8, 18] }   # re-entry

bonuses:
  - { on: basic_3, requires_path: [basic_1, basic_2], damage_multiplier: 1.5 }

(The name fields keep the original Korean display names — 기본1 is Basic 1, 강공격 the heavy attack, 띄우기 the launcher, 회피 the dodge.)

This spec serves two readers at once. A human reads window: [6, 24] and understands "the heavy attack starts being accepted midway through Basic 2"; a machine parses the same line and uses it for the graph drawing and the path enumeration. One source yields both human understanding and machine verification.

The frame values above (21, 24, [6, 24]) are not measurements — they are example values I constructed for this chapter's explanation (unverified). In a real project, these values come from the lengths of the montages your animators build and the notify timings in the build. When you first write the spec, you enter the designer's intended values; once a build exists, you capture it and correct the spec with measured values — that correction loop is covered in 4.4.


4.3.3 Worked Transcript — From Natural Language to Spec

I give the rules B drew on the whiteboard to the AI in natural language and have it convert them into the spec YAML. No summarizing — the full prompt, Claude's raw output, and the human verification/rejection are reproduced exactly as they happened.

The Prompt (Full Text)

다음은 전사 캐릭터의 콤보 룰이다. 이걸 콤보 명세 YAML로 변환해라.

규칙:
- 기본 공격은 1, 2, 3타로 이어진다. 가벼운 공격 키.
- 기본2 도중에 강공격 키를 누르면 강공격으로 빠질 수 있다.
- 강공격 다음에 강공격을 한 번 더 누르면 적을 띄운다(띄우기).
- 강공격은 언제든 회피로 캔슬 가능.
- 기본3도 회피로 캔슬 가능.
- 회피 직후에는 가벼운 공격으로 기본1을 다시 시작할 수 있다.
- 기본3은 기본1과 기본2를 모두 거쳤을 때만 데미지가 1.5배.

명세 형식:
- nodes(id, name, duration_frames), edges(from, to, input, window, type),
  bonuses(on, requires_path, damage_multiplier)
- 캔슬 엣지는 type: cancel로 표시.
- 프레임 값은 내가 안 준 건 너가 정하지 말고 TODO로 비워 둬라.

(The prompt, kept in the original Korean above, says: "These are the combo rules for a warrior character. Convert them into a combo spec YAML." It then lists the rules — the 1-2-3 light-attack chain; a heavy-attack branch available during Basic 2; heavy followed by heavy launches the enemy (launcher); the heavy attack can always be canceled into dodge; Basic 3 can also be canceled into dodge; right after a dodge, a light attack restarts Basic 1; Basic 3 deals 1.5× damage only if both Basic 1 and Basic 2 were passed through — defines the spec format with cancel edges marked type: cancel, and ends with one constraint: for any frame value I didn't provide, don't decide it yourself; leave it as TODO.)

That last line is the crux. I explicitly told it not to invent frames I hadn't given. It enforces rule 3 of the v3 standards (numeric honesty) at the prompt level.

Claude's Raw Output (Unedited)

명세로 변환했습니다. 프레임 값은 제공되지 않아 TODO로 남겼습니다.

```yaml
character: warrior
combo_id: basic_chain
nodes:
  - { id: basic_1, name: 기본1, duration_frames: TODO }
  - { id: basic_2, name: 기본2, duration_frames: TODO }
  - { id: basic_3, name: 기본3, duration_frames: TODO }
  - { id: heavy,   name: 강공격, duration_frames: TODO }
  - { id: launch,  name: 띄우기, duration_frames: TODO }
  - { id: dodge,   name: 회피,   duration_frames: TODO }
edges:
  - { from: basic_1, to: basic_2, input: light, window: TODO }
  - { from: basic_2, to: basic_3, input: light, window: TODO }
  - { from: basic_2, to: heavy,   input: heavy, window: TODO }
  - { from: heavy,   to: launch,  input: heavy, window: TODO }
  - { from: heavy,   to: dodge,   input: dodge, window: TODO, type: cancel }
  - { from: basic_3, to: dodge,   input: dodge, window: TODO, type: cancel }
  - { from: dodge,   to: basic_1, input: light, window: TODO }
bonuses:
  - { on: basic_3, requires_path: [basic_1, basic_2], damage_multiplier: 1.5 }
```

확인이 필요한 부분:
1. "회피 직후 기본1 재시작"을 dodge -> basic_1 엣지로 넣었는데, 이게
   회피의 후딜 중에 입력받는 건지 회피가 완전히 끝난 뒤인지 불명확합니다.
2. 띄우기(launch) 다음에 이어지는 공중 콤보가 있는지 규칙에 없습니다.
   띄우기가 콤보의 끝인지 확인이 필요합니다.
3. 회피를 노드로 둘지 엣지(캔슬 동작)로만 둘지 모호합니다. 일단 노드로
   뒀습니다 — 회피 자체가 지속 시간을 갖는 액션이라고 봤습니다.

(In the raw output above, Claude returns the YAML with every frame value left as TODO and closes with three questions: 1 — it modeled "restart Basic 1 right after the dodge" as a dodge -> basic_1 edge, but it's unclear whether input is accepted during the dodge's recovery or only after the dodge fully ends; 2 — the rules say nothing about an aerial combo after the launcher, so is the launcher the end of the combo?; 3 — it's ambiguous whether dodge should be a node or only an edge (a cancel action); it chose node for now, on the grounds that a dodge is itself an action with a duration.)

Let's evaluate. The good: it didn't invent a single frame — everything is TODO. It honored the prompt's constraint. And the three questions at the end are all genuinely ambiguous points — especially question 1 (re-entry input timing), which is exactly the hole B couldn't answer on the spot in the meeting room. The AI found the hole before the humans did.

The debatable: making dodge a node, in question 3, is a judgment call. A dodge is both a "cancel action" and an "action with a duration," so both readings are valid. The AI picking one side and then reporting the ambiguity is honest, but this is a design decision — a human has to make it.

Human Verification/Rejection

I answer the three questions and reject part of the output.

And one thing gets rejected. The AI didn't put type: cancel on the dodge → basic_1 edge — and in nature it is a cancel, since it cuts the dodge's recovery to enter. But here I treat it as "normal entry after the dodge" and keep it a regular edge, because for this character the distinction between a recovery cancel and a normal link makes no difference in game feel. A case of a human overriding the AI's classification with a domain judgment.

The Follow-Up Request

좋다. 다음을 반영해 최종 명세를 다시 내라:
- dodge에 cancels_recovery: true 추가.
- dodge -> basic_1 엣지의 window는 [8, 18].
- 나머지 프레임은 여전히 내가 안 줬으니 TODO 유지. 단 위 그래프 예시값
  (basic_1=21, basic_2=24, basic_3=30, heavy=33, launch=28, dodge=18)을
  쓸 거니까 그 값으로 채워라. 이건 미검증 예시값이라고 주석으로 박아라.

(The follow-up, in Korean above, says: good — now produce the final spec with cancels_recovery: true added to dodge; window: [8, 18] on the dodge -> basic_1 edge; the remaining frames are still ones I didn't give, so they stay TODO in principle — but since we'll use the example values from the graph above (basic_1=21, basic_2=24, basic_3=30, heavy=33, launch=28, dodge=18), fill them in with those values and pin a comment marking them as unverified example values.)

The result of this follow-up is the YAML in 4.3.2. It wasn't done in one shot: prompt → raw output → verification/rejection → follow-up. That cycle is what gives the spec its credibility. The AI marks what's ambiguous, and the human decides with domain knowledge — neither alone is enough.


4.3.4 Enumerate the Paths Automatically

Since the spec is a graph, enumerating combo paths becomes a graph traversal problem: a depth-first search (DFS) that finds every path from the start node to an end node (or the finisher). A human can't do this in their head; code does it in an instant.

Inside my team's isolated workspace 95_BattleTF lives a small script responsible for this enumeration. It reads the spec YAML, extracts every path, and verifies that each path is legal under the rules (that every edge exists). The core logic looks like this.

# 95_BattleTF/enumerate_paths.py (excerpt)
import yaml

def load_graph(path):
    spec = yaml.safe_load(open(path, encoding="utf-8"))
    adj = {}
    for e in spec["edges"]:
        adj.setdefault(e["from"], []).append(e)
    return spec, adj

def enumerate_paths(adj, start, max_depth=8):
    results = []
    def dfs(node, path, edges):
        # end node (no outgoing edges) or depth limit: finalize the path
        outs = adj.get(node, [])
        if not outs or len(path) >= max_depth:
            results.append((list(path), list(edges)))
            return
        for e in outs:
            if e["to"] in path:        # cycle guard: each node once per path
                results.append((list(path), list(edges)))
                continue
            dfs(e["to"], path + [e["to"]], edges + [e])
    dfs(start, [start], [])
    return results

Run it starting from basic_1, and out pour the paths you could never fully unfold by hand. A sample:

# Path Note
1 Basic 1 → Basic 2 → Basic 3 Textbook 3-hit chain; finisher bonus satisfied
2 Basic 1 → Basic 2 → Heavy → Launcher Branch combo
3 Basic 1 → Basic 2 → Heavy → Dodge → Basic 1 → … Cycle entry
4 Basic 1 → Basic 2 → Basic 3 → Dodge → Basic 1 → … Reset after the finisher

Paths 3 and 4 are the important ones. Because of the dodge re-entry edge, the combo cycles. These cyclic paths are exactly what people failed to see on the whiteboard. Without the cycle guard in the DFS (each node at most once per path), the enumeration falls into an infinite loop — a trap I actually hit the first time I ran the code and it hung. If the graph has cycles, the enumerator must have a guard.

The enumeration stage produces two artifacts. First, the list of every path that is legal under the rules. Second, rule contradiction detection — if the spec has a dodge → basic_1 edge but the dodge node definition is missing, the enumerator flags it as "an edge pointing to an undefined node." That dangling reference is the single most common mistake when specs are written by hand.


4.3.5 Run the Enumerated Paths Through Simulation

A path list alone doesn't tell you which path is too strong. Each path has to go into a DPS simulator. My team's simulate_dps plays that role — it takes a path (a node sequence), each node's damage and frames, and the bonus rules; computes total damage and total elapsed frames; and produces damage per second (DPS).

# 95_BattleTF/simulate_dps.py (excerpt, assumes 60fps)
def simulate(path_nodes, node_dmg, node_frames, bonuses):
    total_dmg = 0
    total_frames = 0
    visited = []
    for nid in path_nodes:
        dmg = node_dmg.get(nid, 0)
        # bonus: apply the multiplier once every requires_path node was visited
        for b in bonuses:
            if b["on"] == nid and all(r in visited for r in b["requires_path"]):
                dmg *= b["damage_multiplier"]
        total_dmg += dmg
        total_frames += node_frames[nid]
        visited.append(nid)
    seconds = total_frames / 60.0
    return {"dmg": total_dmg, "frames": total_frames,
            "dps": round(total_dmg / seconds, 1) if seconds else 0}

Pipe the entire enumeration output from 4.3.4 through this, and per-path DPS falls out as a table. Below is a run with example node damage values (basic hit 100, heavy attack 180, launcher 140 — all unverified values constructed for illustration).

Path Total Damage Total Frames DPS
Basic 1→Basic 2→Basic 3 (finisher ×1.5) 100+100+150 = 350 75 280.0
Basic 1→Basic 2→Heavy→Launcher 100+100+180+140 = 520 106 294.3
Basic 1→Basic 2→Basic 3→Dodge→Basic 1 350+0+100 = 450 144 187.5

This table changes the discussion. The intuition "doesn't the heavy branch look stronger than the textbook 3-hit chain?" becomes a number: "heavy path DPS 294 vs. textbook 280 — a 5% edge." If the 5% edge is intended, it passes; if not, you lengthen the heavy attack's frames to bring the DPS down. And you make that call before a build exists, at the spec stage.

The whole workflow on one page:

flowchart LR
    A["Natural-language rules
(teammate_b's whiteboard)"] --> B["Prompt → AI"] B --> C{"Raw spec
with TODOs and questions"} C -->|human verification/rejection| D["Final spec YAML
warrior_basic_chain.yaml"] D --> E["enumerate_paths.py
DFS enumeration of all paths"] E --> F["Rule contradiction detection
dangling edges, infinite cycles"] E --> G["simulate_dps.py
per-path DPS calculation"] G --> H["Balance judgment
DPS gaps between paths"] F --> D H -->|frame adjustments| D classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; class E,F,G code; class B,C ai; class H human; class A,D data;

The spec (D) sits at the center, and enumeration (E), contradiction detection (F), and simulation (G) all branch off of it. When a contradiction is caught or the balance is off, you go back to the spec and fix it. The whiteboard had no such loop — which is why the whiteboard's combo turned out to be wrong only after it went into a build.


4.3.6 Cancels and Input Queues — Two Levers That Widen and Narrow the Paths

So far we've looked at the combo graph and path enumeration. Cancels and the input queue are the two levers that tune this graph — and they point in opposite directions.

Cancels add edges. Every cancel rule you add puts another edge on the graph, and the number of enumerated paths multiplies. So cancels are not "the more generous, the better." The looser the cancels, the more the paths explode, and the higher the odds that an unintended strong path (like the cyclic path in the previous section) slips in. This is why the fighting-game tradition keeps cancels strict and why action RPGs keep them generous — the genre decides how many paths to allow. There is no absolute correct window value.

When you handle cancels, always separate them explicitly. Leave them as one blanket "cancel into anything" rule, and the enumerator generates cancel edges between every pair of nodes — the paths grow out of control.

Cancel Type Spec Representation Effect on Paths
Action cancel specific node → specific node, type: cancel Adds only selected branches
Dodge cancel many nodes → dodge, window: [0, dur] An escape hatch from almost every node
Guard cancel many nodes → guard Entry into guard, usually limited to recovery
No cancel no outgoing cancel edges Committed to the end once started (super armor)

The input queue doesn't narrow the paths — it makes them actually walkable. (This is the system fighting games call input buffering.) Without the queue, the player would have to hit each edge's input window (say, [12, 24]) with frame-level precision. Human reaction times make that nearly impossible. The queue stores an input pressed before the window into a buffer, then fires it automatically the moment the window opens. In other words, the queue doesn't change the graph's paths; it puts shoes on the player so a human can walk the graph.

input_queue:
  window_start_ratio: 0.5   # buffer the next input from 50% of action progress
  expire_frames: 10         # how long a buffered input stays valid
  priority: latest          # on simultaneous inputs, the last one wins

The balance of the three parameters is the crux. If window_start_ratio is too small, inputs from early in the action get buffered too, and unintended follow-up actions pop out. If expire_frames is too short, the queue becomes meaningless and you're back to demanding precision; too long, and an input pressed ages ago fires late — the "why did my character suddenly move?" incident. A recommended starting point is expire_frames 5–15 and window_start_ratio around 0.5 — but that is a starting line to adjust for genre and character weight, not a correct answer.

One operational note: do not give every character its own input-queue parameters. Keep one global default and override only the characters whose weight is genuinely different (giant boss types and the like). Manage twenty characters' queue values separately and you can no longer tell which differences are intended and which are mistakes.

One more thing to pin down. So far we've treated edges as "exists / doesn't exist" — but even when an edge exists, how the action actually transitions from that node to the next is a separate decision. There are three core options for linking combo actions. They are beyond this book's depth, so I won't cover them in code, but leaving them unmentioned would make the spec only half the picture.

Given the same edge, these three produce opposite game feel. At the spec stage, you usually fix only the edge's existence and its window, and decide the transition method (blending vs. frame skip) in the build, together with the animators. Still, reserving a one-line field in the spec — transition: blend / transition: skip — means nobody has to ask "how did we decide this edge transitions, again?" at the build stage. The transition method is the combo graph's hidden third axis.


4.3.7 Build Verification Is a Separate Job (a 4.4 Preview)

Every verification so far happened on the spec. The path enumeration, the DPS simulation, the contradiction detection — all of it targeted the YAML. But the spec's frame values are the designer's intent, not measurements from the build. The actual length of the montage the animator made, the frame where the build's notify actually fires, the window where the input queue actually operates in the engine — those have to be captured and measured from the build.

Automatically extracting the five signals (hit-start frame, recovery, cancel window, input queue, hitstop) from build footage is hard to implement; realistically, the most trustworthy source is in-game telemetry — have the build log "this action accepted this input at this frame," then check that log against the spec (see 4.4 for a comparison of capture methods). That comparison loop is the subject of 4.4. When a window that was [12, 24] in the spec measures [14, 26] in the build, you correct the spec to the build's measured values.


4.3.8 Common Mistakes and How to Avoid Them

Mistake Why It's Dangerous How to Avoid It
Writing combos as a straight-line table Branches and cycles hide between rows and get missed Spec as a graph (nodes + edges); don't draw by hand
Lumping cancels into "cancel anything" Enumerated paths explode; strong paths slip in Specify action/dodge/guard cancels separately
No cycle guard in the enumerator Infinite loop at the dodge re-entry Guard: each node once per path
Mistaking spec frames for measurements Intent values and build values differ Mark values as intent; correct via build capture (4.4)
Per-character input-queue values Can't tell intended differences from mistakes Global default + override only a few
Trusting AI-filled frames as-is Made-up numbers enter the spec Enforce "values I didn't give = TODO" in the prompt

Key Takeaways


Try It Yourself — A Mini Pipeline for Combo Path Enumeration

This is the minimum procedure you can run along with by hand. All you need is Python and pyyaml.

setup. Create one working folder and put the spec file and the two scripts inside.

combo-mini/
  warrior_basic_chain.yaml   # the spec from 4.3.2
  enumerate_paths.py         # the DFS enumerator from 4.3.4
  simulate_dps.py            # the simulator from 4.3.5

After pip install pyyaml, paste the YAML from 4.3.2 into the spec file as-is.

prompt. Leave the natural-language-to-spec step to the AI. Use the prompt from 4.3.3 as-is, and make sure to include the final constraint.

프레임 값은 내가 안 준 건 너가 정하지 말고 TODO로 비워 둬라.
캔슬 엣지는 type: cancel로 표시하고, 모호한 부분은 질문으로 따로 빼라.

(The two lines say: for frame values I didn't give, don't decide them yourself — leave them as TODO; mark cancel edges with type: cancel, and pull anything ambiguous out as separate questions.)

These two lines block the AI from fabricating numbers and making arbitrary calls. A human then fills in the TODOs and the question list in the resulting spec.

verify. Once the spec is complete, verify it twice.

python enumerate_paths.py warrior_basic_chain.yaml   # all paths + contradictions
python simulate_dps.py    warrior_basic_chain.yaml   # per-path DPS table

In the enumeration output, check (1) that cyclic paths don't grow without bound, and (2) that there are no dangling edges pointing at undefined nodes. In the DPS output, check that the gaps between paths fall within your intended range. If a gap is too large, fix the frames/damage in the spec and run it again.

Solo Scale-Down

If you don't have time to build the tools, finish the spec writing and the path enumeration in a single AI conversation. Give it the natural-language rules and pack this into one prompt: "Convert these into a combo spec YAML, then list every possible path from the start node depth-first, cutting each cycle after one pass. Flag any edge that points to an undefined node." The AI does the spec conversion, the enumeration, and the contradiction detection in one go. For the DPS simulation, hand it the node damage as a table and follow up with "compute each path's total damage and frames and give me a table." You lose precision, but you see much farther than a whiteboard. The core doesn't change — don't unfold combos in your head; have them enumerated, and look.

4.4 AI-Assisted Combat Simulation and Verification

Build #234 from the combat task force (TF) has just come in. It's the first build to touch the new skill skill_thunder. The spec sheet says hit timing: 150ms. I press the input. My fingertips tell me: late. Definitely late. I call over teammate A in the next seat. "Doesn't this look a bit floaty to you?" Teammate A tries it a couple of times. "Hmm... maybe, I guess." Neither of us is sure. The spec says 150; my hands insist it's around 200. Who is right? Fingertips versus paper. In the next build, someone will again say "feels fine to me," and on that one remark another build will drift by.

The goal of this chapter is to end that fight. When the fingertips say 200, show in numbers whether it really is 200. And before a build comes in, know from the spec alone that "this skill's DPS is 30% over target." Back when I was shaping combat from the early days of a AAA MMORPG with 200 people on it, the helplessness when feel and numbers diverged was exactly the same. The only thing that has changed is that now there is a tool at hand to settle that divergence with numbers.

If 4.2 and 4.3 covered how to write a combat spec, 4.4 covers whether that spec works as intended. Verification has two axes. One is simulation — verifying by computation alone, with no build; the other is capture analysis — pulling measured values out of the actual build. When the two axes are tied into one cycle, the turnaround of combat design shrinks from days to hours.

To state the conclusion up front: the heart of this chapter is putting the number written in the spec (150ms) and the number measured from the build (220ms) side by side and reading the gap (4.4.5). The earlier sections (the simulator, combo enumeration) can be read as the preparation that makes that comparison possible.


4.4.1 The Cost of Waiting for a Build

What has to happen for a combat designer to verify one new skill?

The designer writes the spec. A programmer enters the data, an artist attaches motion and effects, a build runs, QA does a pass, and only then does the designer get hands on it. Two days if fast, usually three or four. When "the DPS is too high" is discovered at the end of that cycle, the discovery becomes an order to go back to the start. Another three or four days.

Simulation is a tool that gives an answer at the first step of this cycle. Compute from the spec alone. If the answer is bad, fix the spec and compute again. Before entering the expensive build stage, the spec itself gets filtered once. It's like putting a model car into a wind tunnel first. Before the real car ever touches the road, a suspect design fails on the desk.

Of course, the wind tunnel doesn't predict the road 100%. That's why the second axis, capture analysis, is needed. If simulation is the ideal answer, capture is the answer of what actually happened in the build. Putting the two side by side and reading the difference — that is this entire chapter.


4.4.2 simulate_dps — A Runnable Simulator

Abstract pseudocode verifies nothing. So from the start, I build code that runs. Below is the core skeleton of simulate_dps.py, the one I use in the combat TF, reconstructed for this book with company data stripped out. It runs on the Python standard library alone, with no dependencies (the full file is in "Try It Yourself").

The input is simple. A skill is a dataclass with damage, cast_sec (cast occupancy time), cooldown_sec, and resource_cost; a character has a resource pool, regen per second, a skill list, and a priority rotation. The engine running on top is a single greedy rule: "at every moment, use the highest-priority skill that's available." The point is to capture an ideal upper bound — neither smarter nor dumber than a real player. Excerpting just the spine:

# Build the timeline in 0.05s ticks. When not casting, use the first available skill in priority order.
while t < duration_sec:
    resource = min(char.max_resource, resource + char.resource_regen * tick)
    for name in cooldowns:
        cooldowns[name] = max(0.0, cooldowns[name] - tick)
    if t >= busy_until:                     # wait while the cast motion runs
        for name in char.rotation:          # in priority order
            s = skill_by_name[name]
            if cooldowns[name] <= 0 and resource >= s.resource_cost:
                total_damage += s.damage
                resource -= s.resource_cost
                cooldowns[name] = s.cooldown_sec
                busy_until = t + s.cast_sec  # no next skill until this time
                break
    t += tick
# …(for the dataclass definitions, warrior input, and output loop, see the full code in "Try It Yourself")

Running warrior for 20 seconds with skill_thunder (damage 420, cast 0.9s, cooldown 6s) at priority 1 and skill_dash and basic_1 behind it (python simulate_dps.py):

평균 DPS: 261.0
  t=  0.0s  skill_thunder  자원=60
  t=  0.9s  skill_dash     자원=47
  t=  1.3s  basic_1        자원=50
  t=  1.6s  basic_1        자원=53
  t=  1.9s  basic_1        자원=55
  ...

What does this value mean? It means that without a build, within one second, I know the warrior's ideal DPS ceiling is about 261. If the target DPS was 180, this spec is +45% over. No need to wait for a build — adjust damage or cooldown_sec right now.

I'll state the limits honestly too. This simulator does not model player input mistakes, gaps caused by movement and dodging, or enemy interference. So its number always comes out higher than the actual build. That is not a bug; it is the very definition of the simulator as an upper bound. The gap to reality gets filled by capture in 4.4.5.

A note on using AI. I wrote the skeleton above myself, but when attaching a new resource model (say, a rage gauge that fills when taking damage), I ask Claude while quoting the existing code: "Add a rule to this simulate_dps — +5 rage on getting hit. Inside the tick loop, as a separate variable from the existing resource regen." Ask it to generate a whole simulator from a blank page and you get unverifiable code. The human holds the spine; the AI grows the branches.


4.4.3 Automatic Combo Path Enumeration

A single DPS number is not enough. To verify "which combo is the intended main combo," you have to lay out every possible path. Draw the tree by hand and your head bursts at seven or eight nodes. Exhaustively unfolding paths is something a machine does overwhelmingly better than a person — provided you make it produce output a person can trace back through.

The following code takes a combo graph, enumerates all paths, and sorts them by DPS. combo_graph is an adjacency list of "which action can cancel into which action," extracted directly from the state machine spec of 4.3. The core is the all_paths generator, which recursively unfolds every path to its dead end.

# enumerate_combos.py — unfold every path in the combo graph and sort by DPS
combo_graph = {"start": ["A"], "A": ["B", "D"], "B": ["C", "E"], "D": ["C"], "C": [], "E": []}
action_stats = {  # (damage, duration in seconds)
    "A": (300, 0.8), "B": (450, 1.0), "C": (450, 1.2), "D": (600, 1.4), "E": (200, 0.6),
}

def all_paths(node="start", path=None):
    path = (path or [])
    nexts = combo_graph.get(node, [])
    if not nexts:                       # dead end = a completed combo
        yield [n for n in path if n in action_stats]
        return
    for nxt in nexts:
        yield from all_paths(nxt, path + [nxt])

results = []
for p in all_paths():
    dmg = sum(action_stats[a][0] for a in p)
    dur = sum(action_stats[a][1] for a in p)
    results.append((p, dmg, round(dur, 1), round(dmg / dur, 1)))

for p, dmg, dur, dps in sorted(results, key=lambda r: -r[3]):
    print(f"{' → '.join(p):<18} {dmg:>5} dmg  {dur:>4}s  DPS {dps}")

The output:

A → D → C            1350    3.4s  DPS 397.1
A → B → C            1200    3.0s  DPS 400.0
A → B → E             950    2.4s  DPS 395.8

The signal a designer should read here is not simply which path came first. The three paths' DPS values sit nearly on top of each other, at 396–400 — and that is a signal that "every combo is about equally efficient, so the main combo has no identity." If the intent was "A→D→C should be the high-risk, high-reward main combo," then raise D's damage or shorten its time to lift that DPS a clear step above the rest. Time to go back to the spec.

Where this automatic enumeration replaces hand calculation, even when the combo nodes grow to 20, a person only has to read the sorted table.


4.4.4 Build Capture Analysis — The Realistic Method Is Telemetry

Now the second axis: measuring what actually happened in the build. The picture that comes to mind first is "an AI watches the build footage and analyzes it automatically," but here I'll split the options honestly. There are three ways to get measurements, and their cost and accuracy differ greatly.

A. Direct video analysis Extract input · motion · VFX from screen pixels Accuracy: low–medium Difficulty: very high Frame error ±1~2 For research · demos B. Off-the-shelf vision API Send captured frames to an external vision service Accuracy: medium Difficulty: medium IP leak risk Check company policy first C. In-game telemetry The engine logs events with timestamps Accuracy: high Difficulty: low–medium Uses the engine clock itself The realistic choice

Option A (direct video pixel analysis) sounds attractive: automatically extract five signals from the screen — the input display, character motion changes, the first frame of an effect, the audio waveform, the damage-number UI. But build it for real and an error of ±1–2 frames is the baseline, thanks to frame compression noise, UI occlusion, and motion blur. At 60fps, one frame is about 16.7ms. For verification that argues hit timing in milliseconds, ±33ms of noise is fatal. The implementation difficulty is very high, and the accuracy doesn't repay the effort.

So the realistic answer is option C: in-game telemetry logs. The engine already knows, internally and exactly, the input time, the animation notify time, the VFX spawn time, and the damage application time. Instead of inferring those times from pixels, make the engine print them as one-line logs. Rather than reconstructing 100ms from pixels, you take down verbatim the 100ms the engine already knows.

// One line added to the combat action handling code (UE C++ pseudo-example)
// Call the same logger at input receipt / damage application
CombatTelemetry::Log("input",  SkillName, GetWorld()->GetTimeSeconds());
CombatTelemetry::Log("hit",    SkillName, GetWorld()->GetTimeSeconds());

The logger drops one JSON Lines entry per line.

{"event":"input","skill":"skill_thunder","t":12.340}
{"event":"hit",  "skill":"skill_thunder","t":12.560}
{"event":"input","skill":"basic_3","t":14.100}
{"event":"hit",  "skill":"basic_3","t":14.166}

The time difference between input and hit is the measured hit timing. 12.560 − 12.340 = 0.220 seconds = 220ms. Not the ±33ms of pixel analysis — the engine's own clock, as is. Pulling this log and checking it against the spec is the next section.


4.4.5 Worked Transcript: Spec 150ms vs. Measured 220ms

Now bring the two axes to one table. The spec promised 150ms; telemetry measured 220ms. +70ms. The fingertips were right. I'll follow, end to end with the AI, the process of narrowing down the cause of this gap. No summarizing — the scene where the AI grabs the wrong lead at first is included as is.

Prompt

The spec says hit timing 150ms; the measurement says 220ms. I paste the log below and ask at the length a person would actually type. A short request with clear intent beats elaborate formality — one line demanding the verification method and one line banning assertions are the spine of this prompt.

skill_thunder hit timing is spec'd at 150ms but measures 220ms, +70ms over. Using the telemetry below, break down the possible causes and state for each one which log confirms or rules it out. No assertions — if you're not sure, mark it 'unverified'. // The two key intents are "include the verification method" and "no assertions." Drop the rest of the formality and the result is the same.

{"event":"input","skill":"skill_thunder","t":12.340}
{"event":"anim_start","skill":"skill_thunder","clip":"cast_thunder","t":12.340}
{"event":"anim_notify","skill":"skill_thunder","notify":"HitWindow","t":12.560}
{"event":"hit","skill":"skill_thunder","t":12.560}

Per the spec, the HitWindow notify of cast_thunder is at the 0.150s mark.

Claude's Raw Output (Round 1 — Before Verification, Wrong Lead Included)

I estimate the following four causes for the +70ms overshoot.

  1. Network latency. The input may have picked up 70ms on a round trip to the server. If this is a multiplayer build, check the RTT.
  2. Frame rate drop. If the build runs at 30fps, one frame is 33ms, so notify timings get rounded to frame boundaries and the delay accumulates.
  3. Cast motion length change. The cast_thunder clip may have become longer than spec'd, pushing the HitWindow notify back.
  4. Input buffer delay. The input queue may have been processed one tick late.

The most likely cause is #1, network latency. 70ms is a typical RTT range.

Human Verification and Rejection

Stop here and reread the log. The AI's top-ranked guess (network) is already ruled out by this very log. input and anim_start are stamped at exactly the same time, 12.340. The motion started the instant the input arrived; there is no room for a network round trip. #1 is wrong.

2 (frame rate) is weak too. At 30fps you would see jitter quantized in 33ms steps, but input→hit is a clean, exact 0.220. That smells not of frame-boundary rounding but of a fixed position inside the clip.

The decisive clue is elsewhere. anim_notify fires at +0.220 relative to anim_start. The spec says HitWindow should sit at 0.150 after clip start, but in the actual clip it is fixed at the 0.220 mark. In other words, either the clip itself was authored differently from the spec, or the notify position was moved from 0.150 to 0.220. The cause is neither network nor frame rate — it is #3, the motion/notify data. The AI did include the right answer among its four candidates, but ranked them upside down.

Follow-Up Prompt

I fill in the exclusion grounds by hand and throw it back. No need for a long explanation — just nail down, briefly, what was ruled out and which facts remain.

Network and frame rate ruled out. (Grounds: input = anim_start at the identical time, so no input→motion delay / input→hit is a constant 0.220 with no frame noise.) Remaining fact: anim_notify is at anim_start +0.220; spec says +0.150. Narrow the cause from this fact alone, and state what to request from the programmer and the animator in units of actions. // The key is "in units of actions" — end not with guesses but with who gets asked for what.

Claude's Raw Output (Round 2)

Given the stated grounds, the cause narrows to the HitWindow notify in animation clip cast_thunder being placed at 0.220s instead of 0.150s. Two branches remain.

Checks that separate the two branches: 1. To the animator: request cast_thunder's HitWindow notify frame number and the clip's total frame count, plus the change history compared against the values at spec-writing time. If the marker was moved, it's (b). 2. To the programmer: confirm whether skill_thunder's montage play rate is 1.0, and whether the character's attack speed stat multiplies into this clip. If the play rate is 0.68 (=0.150/0.220), it's (a). 3. Measure the anim_start→anim_notify interval once more at a different character attack speed. If the interval changes with attack speed, it's (a); if it stays invariant, it's (b).

Recommended action: if confirmed as (b), move the notify back to 0.150 or update the spec to 0.220 (whichever was intended). If (a), inspect the play-rate application logic.

This output is worth adopting. It does not assert a guess; it splits the question into branches (a)/(b), and how to separate the branches with data and who to ask for what are written in units of actions. Check #3 in particular (re-measuring at a changed attack speed) is a decisive fork that a person easily misses. Take this report into the meeting as is, and the meeting stops being a debate over "what could the cause be" and becomes "confirm (a) or (b) within 30 minutes and pick the action."

What this worked transcript shows, above all, is that the AI does not hand you the right answer from the start. The first output was a wrong lead that ranked network first. Only when human verification — reading the log and ruling out candidates — entered the loop did the analysis converge on the answer. The AI spreads the candidates wide; the human narrows them. That division of labor is the methodology of all of 4.4.


4.4.6 Two Axes, One Cycle — The Verification Loop

If simulation (4.4.2–4.4.3) and capture analysis (4.4.4–4.4.5) run separately, you get half the value. Tied together, they form the following loop.

flowchart TD
    A["Write the spec
(designer, 4.2·4.3 templates)"] --> B["Simulation
simulate_dps · combo enumeration
(automated, 1 second)"] B -->|"DPS/combo anomaly"| A B -->|"Spec passes"| C["Build
(programmers & artists, days)"] C --> D["Collect telemetry logs
(input · hit · anim_notify)"] D --> E["Auto-compare spec vs. measurement
(detects gaps like 150ms vs 220ms)"] E -->|"Items over threshold"| F["AI cause-hypothesis report
(AI widens candidates → human narrows)"] F --> G["Designer review → decide action
(update spec vs. fix build)"] G --> A E -->|"All items match"| H["On to the next skill"] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; class B,E code; class F ai; class A,G human; class D data; class H pass;

Because simulation gives its answer before the build, any spec that enters a build has already been filtered once. So a problem found after the build narrows in character: not "the spec was wrong" but "the spec and the implementation diverged." The +70ms of 4.4.5 was exactly the latter — the spec's 150 was reasonable; the implementation drifted to 220. This distinction removes blame disputes from meetings.

This loop does not cover all combat content, though. Areas where the feel is the content — a main boss's signature set piece, say — do not reduce to a DPS number. Simulation is strong on numbers-driven content; for presentation, the answer is still human eyes reviewing the footage. The loop runs in the river of numbers; the river of presentation flows on its own.


4.4.7 What Six Months of Operation Showed

These are six months of measurements from the combat TF of an MMORPG I ran (Project A, which targeted refgame-style controls). The figures below are actual measurements culled from the TF's internal records; I'll note that build counts and times are rounded to cycle units (measured in cycles and half-days, not exact minutes).

Item Before After
Verification cycle for a new skill 3–4 builds on average 1–2 builds on average
Verifying 100 skills in a build Half a day (manual) 30 minutes (automated telemetry comparison)
Balance meeting 2 hours (subjective debate) 30 minutes (data-driven)
Defects found just before a build 5–8 per build on average 1–2 per build on average

The change that matters more than the numbers is the character of the meetings. Before, half of every meeting was "this skill is too strong" versus "no, it's fine." Fingertips versus fingertips. After, that time moved to "measured DPS is +12% over target and measured hit timing is +70ms over spec — do we shorten the motion, or cut damage by 10%?" The time spent fighting over what the problem is became time spent deciding how to fix it.

I'll write down the cost of this transition honestly too. Planting the telemetry logger across the combat code took about one to two weeks of initial work, and unifying the spec schema across every character and skill took another quarter on top of that. In the first quarter, I judged it sufficient for just one of the two — simulation or telemetry — to be running. The two only meshed into a single loop from the second quarter on.


4.4.8 Common Mistakes and How to Avoid Them

Five mistakes keep repeating.

First, trusting the simulation value absolutely. simulate_dps's 261 is an upper bound, not a measurement. Without comparing against telemetry, you will always overestimate.

Second, not measuring at all. If you only simulate the spec and never look at the build through telemetry, gaps like the +70ms of 4.4.5 quietly accumulate. Telemetry logging is not an option; it is basic plumbing of combat code.

Third, not bringing the report to the meeting. Even with the automated report sitting right there, nobody looks at it if it isn't on the agenda. Make "this build's telemetry comparison table" a fixed agenda item.

Fourth, adopting AI estimates without verification. As we saw in 4.4.5, the AI's first estimate was a wrong lead. AI is a tool for widening the candidates, not for reaching the conclusion. Never skip the human step of ruling out candidates with logs.

Fifth, letting every skill have a different spec structure. If one skill writes cast_sec, another cast_time, and another castMs, both the simulator and the comparison script break every time. Enforce one common spec schema on every skill — that is the premise on which all of 4.4's tools run.

You don't need to fix all five in the first quarter. Even one or two taking root shortens the cycle visibly. The rest fill in naturally as the loop keeps turning.


4.4.9 Closing Part 4

Across 4.1–4.4 we covered, in order, the coordinates of combat design, look & feel, combos, and simulation. 4.1 covered what a combat designer treats as a measurable object; 4.2, how to measure and tune hit timing, hitstop, and effect synchronization; 4.3, how to write combos, cancels, and input queues as state machines; and 4.4, how to verify — without a build and in the build — that all those specs work as intended.

A combat designer's week after finishing Part 4 changes like this. Monday: simulate_dps auto-verification attaches to the new skill spec. Tuesday: automatic combo path enumeration checks the main combo's identity. Wednesday: when the build comes in, the telemetry comparison table appears automatically. Thursday: 30 minutes of data-driven discussion. Friday: spec revisions for the next cycle. Build cycles drop from 3–4 to 1–2, and meetings shrink to less than half. The abstraction called impact feel moves over to a measurable 220ms.

And that scene from build #234 — to the question "doesn't this look a bit floaty?", the telemetry log now answers in our place: 220ms. The fight of fingertips versus paper is over.

Next, Part 5 is narrative design. We move on to the full-scale application of the NarrativeDocs Layer 0–4 structure introduced in 2.3.


Try It Yourself

setup. 1. Create the full simulate_dps.py below, exactly as is. No dependencies; it runs immediately with python simulate_dps.py and reproduces the "평균 DPS: 261.0" output from 4.4.2.

# simulate_dps.py — compute DPS from the spec alone, no build needed
from dataclasses import dataclass

@dataclass
class Skill:
    name: str
    damage: float          # damage per hit
    cast_sec: float        # cast (motion-occupancy) time, seconds
    cooldown_sec: float    # cooldown, seconds
    resource_cost: float   # resource cost (MP/stamina)

@dataclass
class Character:
    name: str
    max_resource: float
    resource_regen: float  # resource regen per second
    skills: list           # list[Skill]
    rotation: list         # priority order (skill names)

def simulate_dps(char: Character, duration_sec: float, tick=0.05):
    cooldowns = {s.name: 0.0 for s in char.skills}   # remaining cooldowns
    skill_by_name = {s.name: s for s in char.skills}
    resource = char.max_resource
    total_damage = 0.0
    busy_until = 0.0          # when the cast motion ends
    log = []
    t = 0.0
    while t < duration_sec:
        resource = min(char.max_resource, resource + char.resource_regen * tick)
        for name in cooldowns:
            cooldowns[name] = max(0.0, cooldowns[name] - tick)
        if t >= busy_until:   # pick the next skill when not casting
            for name in char.rotation:          # in priority order
                s = skill_by_name[name]
                if cooldowns[name] <= 0 and resource >= s.resource_cost:
                    total_damage += s.damage
                    resource -= s.resource_cost
                    cooldowns[name] = s.cooldown_sec
                    busy_until = t + s.cast_sec
                    log.append((round(t, 2), name, resource))
                    break
        t += tick
    return total_damage / duration_sec, log

if __name__ == "__main__":
    warrior = Character(
        name="warrior", max_resource=100, resource_regen=8,
        skills=[
            Skill("skill_thunder", damage=420, cast_sec=0.9, cooldown_sec=6, resource_cost=40),
            Skill("skill_dash",    damage=180, cast_sec=0.4, cooldown_sec=3, resource_cost=20),
            Skill("basic_1",       damage=60,  cast_sec=0.3, cooldown_sec=0, resource_cost=0),
        ],
        rotation=["skill_thunder", "skill_dash", "basic_1"],
    )
    dps, log = simulate_dps(warrior, duration_sec=20)
    print(f"평균 DPS: {dps:.1f}")
    for t, name, res in log[:8]:
        print(f"  t={t:>5}s  {name:<14} 자원={res:.0f}")
  1. Add one telemetry log line each at the input-reception and damage-application points of your combat engine code (4.4.4). The key is printing the engine clock (GetTimeSeconds and the like) as is.

prompt. Pick one item in your build telemetry that diverges from the spec and query the AI in the format of 4.4.5 — paste the log excerpt, and be sure to include "don't assert guesses; give a verification method for each cause, and mark anything uncertain as unverified."

verify. Do not adopt the AI's first output as is. Read the log yourself, strike out the candidates you can rule out by hand (like the network and frame-rate exclusions in 4.4.5), attach your grounds, and ask again. When the final report ends in units of actions — "estimated cause + who to ask for what" — take it to the meeting.

Solo Scale-Down. If installing the telemetry logger everywhere is too much, log just the two lines — input and hit — for the one skill you want to verify. Run simulate_dps for that one skill's DPS only. Don't deploy the whole toolset; run the loop once around the single most suspicious skill, then expand.


Key Takeaways

Part 5 · Narrative Design

5.1 The NarrativeDocs Layer 0–4 Structure

When I walked into the meeting room, a character's name was circled in red on the whiteboard. An NPC I'll call Kim. In one designer's side quest, he was "the adoptive father who took the protagonist in as a child." In chapter 3 of another designer's main quest, he was "the old comrade who betrayed the protagonist and left." Both documents had been approved a month earlier, and both were already in the build. We had even gotten a quote for voice recording.

It was nobody's fault. Both designers had read the worldbuilding document and consulted the character's lore. The problem was that the same character's lore was scattered across three different files, and nobody could say with certainty which one was "real." The worldbuilding document was a monolithic 70-page Word file; a search turned up Kim in eleven places. There was no way to tell which lines were decisions and which were memos.

What we decided after that meeting was to split NarrativeDocs into five layers. This chapter is the story of those five layers.


5.1.1 Why Start with the Most Abstract Discipline

There is a reason the discipline-by-discipline Layer decomposition opens with narrative.

Narrative is the most abstract discipline. Worldbuilding, emotion, tone — none of it reduces to numbers. Unlike art or systems, there are no clean units like "sprite count" or "damage coefficient." If a discipline this abstract decomposes cleanly into Layers, the more concrete disciplines naturally follow the same pattern. In effect, we test whether the approach works in the hardest place first.

Narrative also has the most interfaces. Characters touch art; quests touch content and level design; dialogue touches UX and localization; rewards touch systems. It crosses more discipline boundaries than anything else, so the value of unifying the Layers shows up immediately.

Finally, natural-language deliverables make up the largest share of its output — which also makes it the discipline where AI assistance does the most work. That said, not every game is narrative-driven. For a casual or arcade genre, this chapter's depth may be overkill. Even so, the skeleton itself — split the monolithic document into Layers and narrow the interfaces — carries over to any discipline as is.


5.1.2 The Five-Drawer Chest

Here is the shape of NarrativeDocs decomposed into five layers. Vision (L0) sits at the top, build/QA (L4) at the bottom, and the passages connecting the layers are deliberately narrow.

L0 Vision — "What is this world?" world_premise · narrative_pillar · tone_manifesto (immutable · ~4.5 pages) Director · Lead pillar · tone (change triggers L1 re-review) L1 System — "How does it work?" faction_system · reputation_model · dialogue_branching_rule · lore_consistency_rule Senior narrative rulebooks · branching policy (affected quests auto-listed) L2 Content — "What happens?" main_quest · side_quest · character_bible · lore_codex (the thickest layer) Multiple designers quest_id · npc_id · dialogue_id (a single column) L3 Data — "The form machines read" quest_table · npc_table · dialogue_id_table · reward_table (0 lines of natural language) Narrative + data sheet change auto-triggers lint L4 Build·QA — "Can it ship?" narrative_qa_checklist · voice_review_log · localization_status Narrative + QA L0 is injected as context every time (anchor for generation & review)

Mapped to folders, it looks like this. Each filename below is not an abstraction invented for this book — it is a file that actually exists in that folder.

NarrativeDocs/
├── Layer0_Vision/
│   ├── world_premise.md          (worldbuilding premise — immutable)
│   ├── narrative_pillar.md       (3 emotional pillars)
│   └── tone_manifesto.md         (tone + forbidden-word list)
├── Layer1_System/
│   ├── faction_system.md
│   ├── reputation_model.md
│   ├── dialogue_branching_rule.md
│   └── lore_consistency_rule.md
├── Layer2_Content/
│   ├── main_quest/               (per chapter)
│   ├── side_quest/
│   ├── character_bible/
│   └── lore_codex/
├── Layer3_Data/
│   ├── quest_table.xlsx
│   ├── npc_table.xlsx
│   ├── dialogue_id_table.xlsx
│   └── reward_table.xlsx
└── Layer4_Build_QA/
    ├── narrative_qa_checklist.md
    ├── voice_review_log.md
    └── localization_status.md

What matters is that no single person owns all five layers. Each layer has a different primary owner, and only the passages between adjacent layers are standardized. The red arrows in the figure above are those passages. The passage between L2 and L3 is drawn deliberately as the narrowest of all (the thick red arrow) — the reason comes later.

As a chest of drawers, it is a five-drawer chest. The first drawer holds the one worldbuilding line that never moves, the second the rulebooks, the third the prose, the fourth the sheets, the fifth the review logs. The passages between the drawers are narrow, and each passage has a change-notification bell mounted on it.


5.1.3 L0 — Kept Small So It Never Changes

L0 does not change. If it changes, the game's identity changes. So it has to be small. Small is what keeps it from changing.

Document Length
world_premise.md 1.5 A4 pages
narrative_pillar.md 1 A4 page (3 emotions)
tone_manifesto.md 2 A4 pages (tone + forbidden-word list)

About 4.5 pages in total. That is the weight of L0. If it gets heavy, changing it becomes frightening; once changing it is frightening, the other layers start routing around L0. Once the routing-around begins, L0 is a dead document.

The actual skeleton of narrative_pillar.md looks like this (content abstracted).

---
title: Narrative Emotional Pillars
layer: L0
status: locked
last_updated: 2026-05-18
---

## 1. Longing for What Was Lost
- The player loses one thing at the end of every chapter.
- What is lost does not come back (flashbacks only).

## 2. The Conflict Between Duty and Freedom
- Every major NPC carries two duties.
- The player's choice can honor only one of them.

## 3. The Weight of a Small Kindness
- A small kindness yields a greater outcome than a grand heroic act.

These three pillars — longing for what was lost, the conflict between duty and freedom, the weight of a small kindness — set the direction for the hundreds of pages beneath them. And status: locked is not a mere label. One of L4's automated checks reads it: if a locked document is modified in a PR, the merge is blocked without the lead narrative designer's approval.


5.1.4 L1 — The Narrative Closest to Code

L1 can change, but changes are expensive. It is the rulebook layer. When one line of a rulebook changes, every piece of content that follows that rule is affected.

Here is the skeleton of faction_system.md.

---
title: Faction System
layer: L1
atoms:
  - faction_relation_matrix
  - faction_membership_rule
  - faction_quest_eligibility
---

## 1. Faction Concept
N factions. Each faction is defined by (ideology, resources, territory).

## 2. Inter-Faction Relations
- relation_matrix.json (-3 hostile ~ +3 allied)
- Relation change triggers: main quest decisions, reputation thresholds

## 3. Player Membership Rules
- Maximum 2 simultaneous memberships (no two hostile factions at once)
- Leaving penalty: reputation -2, allied factions -1

Look at the atoms: list in the frontmatter. These three atom names are not just memos — they are the identifiers tracked by the Part 7 ontology and the Part 11 relation map. When a quest references faction_quest_eligibility, that quest is automatically added to the impact list whenever this rule changes. L1 is the narrative deliverable closest to game code, so we write it paired with a systems designer.


5.1.5 L2 — The Thickest Layer, Where the Prose Lives

L2 is the thickest layer. Main quests, side quests, the character bible, and the lore codex all live here. The real lore for Kim — the "adoptive father vs. traitorous comrade" who collided in that meeting room — now exists only as one file inside character_bible/. That file is the single source of truth, and quests only reference it.

The main quest folder looks like this.

main_quest/
├── chapter_01_awakening/
│   ├── 00_chapter_overview.md
│   ├── 01_quest_a_call_to_arms.md
│   ├── 02_quest_b_first_choice.md
│   └── ...
├── chapter_02_road/
│   └── ...
└── _TEMPLATES/
    └── quest_template.md

Each quest file follows the standard atom format.

---
title: When You Take Up Arms
layer: L2
type: main_quest
atoms:
  - quest_chapter_01_awakening_a
related:
  affects: [reputation_model, faction_relation_matrix]
  derives_from: [narrative_pillar, world_premise]
  requires: [character_kim, faction_alpha]
  part_of: chapter_01_awakening
---

## Progression Steps
1. ...

## Branches
- If option A is chosen: ...
- If option B is chosen: ...

## Rewards (see L3)
- reward_table.xlsx → quest_001 row

The heart of it is the related: block. The moment you write requires: [character_kim], you have declared that this quest pulls Kim's lore from the character bible — it no longer defines Kim anew inside its own file. Narrative prose lives in L2; numeric rewards live in the L3 sheet. If both live in one file, every one-line sheet fix means touching the prose, and then the translation keys drift out of sync.


5.1.6 L3 — Not a Single Line of Natural Language

L3 is sheets and IDs. Not one natural-language sentence goes in.

quest_table.xlsx
| quest_id | chapter | type | unlock_level | reward_xp | reward_gold | dialogue_set_id |
|----------|---------|------|--------------|-----------|-------------|-----------------|
| q_001    | ch01    | main | 1            | 500       | 100         | ds_001          |
| q_002    | ch01    | main | 2            | 800       | 150         | ds_002          |

Even dialogue is referenced only by ID. The lines themselves live separately in dialogue_id_table.xlsx, mapped 1:1 to translation keys. The passage connecting L2 and L3 needs exactly one column: quest_id. That is why this passage alone was the thick red arrow in the earlier figure. Narrowing an interface down to a single column is the core of Layer separation. When the passage is wide, the two sides know too much about each other, and fixing one side breaks the other.


5.1.7 L4 — The Ship Gate

L4 is review and shipping. Every time new content lands, automated and manual checks run. The automated checks are scripts. The four below are the lints actually wired into CI.

Check Tool
Every dialogue_id has a mapping dialogue_lint.py
Every quest_id belongs to a chapter quest_lint.py
Reward totals stay within each chapter's curve range reward_curve_check.py
Forbidden-word occurrences tone_lint.py (based on the L0 tone_manifesto)

Note that tone_lint.py reads L0's tone_manifesto.md directly. The topmost layer (the immutable vision) and the bottom layer (the ship gate) are wired together by automation. If a forbidden word written into the vision turns up in the prose right before release, the build is blocked.

What automation cannot catch, people review.

Check Owner
Fit with the L0 emotional pillars Lead narrative designer
Character voice consistency Character owner + narrative
Localizability Localizer

When the boundary between automated and manual is clear, review time becomes predictable. You can answer the question, "How many days will this chapter's review take?"


5.1.8 Putting AI to Work on the Layers — A Worked Transcript

That is the structure. But where the structure really earns its keep is in making new content. What follows is a real session, reproduced as is, in which I had Claude take a one-line side-quest idea — "a short side quest where Kim hands the player a keepsake from his late wife" — and scatter it across the Layers of a project where the five layers were already in place. Nothing is summarized; the parts that missed and the part I rejected are included. The prompt pinned the L0/L1 context as inviolable and asked for the L2 quest file (with the related: block and about 200 characters of prose), the L3 quest_table row and dialogue_id, and a tone self-check — with instructions to write only L2 first, then stop.

The prompt I sent (in full):

Do not violate the attached L0/L1 context in any way, and decompose this into our Layer structure.

[narrative_pillar.md in full — the 3 emotional pillars]
[tone_manifesto.md excerpt — forbidden words: "fate", "prophecy", "the chosen one"]
[character_bible/character_kim.md excerpt — Kim: the adoptive father who took
 the protagonist in, taciturn, trails off mid-sentence, never names his feelings]

Idea: "A short side quest where Kim hands the player a keepsake from his late wife."

The L2 quest file (with the related block, ~200 characters of prose), the L3 quest_table
row and dialogue_id, plus a tone-violation self-check. Write only L2 first, then stop.

Claude's raw output (L2):

---
title: The Faded Handkerchief
layer: L2
type: side_quest
atoms:
  - quest_side_kim_keepsake
related:
  affects: [reputation_model]
  derives_from: [narrative_pillar]
  requires: [character_kim]
  part_of: chapter_02_road
---

## Progression Steps
1. Kim quietly calls the player over. As usual, he does not state his business right away.
2. He holds out an old handkerchief. "This… I thought it might be better off in your hands."
3. When the player asks what it is, Kim looks out the window instead of answering.
4. (Choice) Accept / Decline.

## Branches
- Accept: Kim's reputation +1. Unlocks 1 handkerchief-related flashback in a later chapter.
- Decline: Kim nods and puts the handkerchief away. No reputation change.
  The handkerchief never appears again.

## Rewards (see L3)
- reward_table.xlsx → quest_side_kim_keepsake row

Claude stopped there, as asked — a quest titled "The Faded Handkerchief," in which Kim quietly calls the player over, wordlessly offers an old handkerchief, looks out the window instead of answering when asked what it is, and an accept/decline branch where declining means the handkerchief never appears again — and appended one line.

"Reflecting L0 pillar 1 (what is lost does not come back), I made the handkerchief disappear forever if the player chooses 'decline.' One thing needs confirmation, though: the dead-wife detail was not in the character_kim.md excerpt. I added it on my own, so it may conflict with the character bible."

Human verification and rejection:

The tone of this output was exactly right. Kim's reticence, his trailing-off sentences, his habit of never naming his feelings — all of it was alive. No forbidden words either. But that final self-check had put its finger on the real problem. Nothing anywhere established that Kim had a dead wife. The AI had made the natural inference from "keepsake" — the Korean word in the idea, 유품, specifically means something left behind by the dead — to a deceased spouse, and that amounts to quietly adding a new fact to the bible. This is precisely the seed of the accident from the meeting room at the start of this chapter.

So I rejected it and sent a follow-up: the tone could stay, but "dead wife" was new lore the bible does not have — Kim's past is confirmed only as far as "he left something behind when he left his hometown" — so do not specify what the keepsake is, blur it to "something left behind long ago," and since the handkerchief is not in the bible either, soften it to a scrap of cloth.

The follow-up request:

The tone is good. But "dead wife" is new lore the bible does not have, so we can't use it.
Kim's past is confirmed only as far as "he left something behind when he left his
hometown." Don't specify what the keepsake is — rewrite it blurred to no more than
"something left behind long ago." The handkerchief isn't in the bible either, so soften it to "a scrap of cloth."

Output after the follow-up (excerpt):

2. He holds out an old scrap of cloth. "There was something… I left behind,
   long ago. Let's just say this is a corner of it."
3. When the player asks what it is, Kim looks out the window instead of answering.

This time I accepted it. Kim now offers an old scrap of cloth — "There was something… I left behind, long ago. Let's just say this is a corner of it." — and again looks out the window when asked. It touches emotional pillar 1 (loss) without nailing down anything new about Kim's past. What was lost stays blurred to the very end.

This one session shows two things. First, when you inject the Layer context precisely (L0 pillars + L1 tone + L2 bible), the AI matches tone remarkably well. Second, when the AI meets a blank, it fills it in plausibly. When that filler is a new fact absent from the bible, catching it is still a human's job. Because the Layer structure's requires: [character_kim] said where to look, the verifier knew immediately which file to check against. Without the structure, it would have meant digging through 70 pages all over again.


5.1.9 Alarm Bells on the Narrow Passages

The real reason for splitting into five layers is to narrow the interfaces. And every narrow interface gets change-detection automation attached to it.

Interface What flows through it
L0 → L1 pillar, tone (a change triggers an L1 rulebook re-review)
L1 → L2 rulebooks and branching policy (a change auto-lists the affected quests)
L2 → L3 quest_id, npc_id, dialogue_id (a single column)
L3 → L4 a sheet change triggers lint automatically

For example, when someone opens a PR on L1's faction_system.md changing "maximum 2 simultaneous memberships" to "maximum 1," the relation map sweeps the L2 quests that reference faction_quest_eligibility, builds the impact list, and attaches it to the PR as an automatic comment. Instead of "book two or three meetings to find out who is affected," the person who changed the rule reads the attached list and finishes with one 30-minute meeting.

The point is this: if you split the Layers but leave the interfaces vague, all you have added is partitions. The essence is not the Layer split itself but the interface automation.


5.1.10 Six Months of Operation: What Changed

These are the measurements after running the five layers for six months. The figures below are based on my team's operating records, but absolute values are carried over only as directions and ratios (author's estimate, unverified). One limitation: "before" comes from memory and "after" from actual measurement, so this is not the same data measured twice.

Item Before the Layer split After the Layer split
New designer onboarding 3 weeks 1 week
Scoping the impact of a rulebook change 2–3 meetings Auto comment + a 30-minute meeting
Producing 1 new main quest chapter 4 weeks 2.5 weeks
Pre-release review of 1 chapter 5 days 2 days
Localization-miss incidents 3–5 per quarter 0–1 per quarter

The clearest drop is localization misses. With dialogue_id tied 1:1 to translation keys in L3 and dialogue_lint.py blocking mapping gaps, the "untranslated line ships in the build" incident all but disappeared. The shorter onboarding mattered too. Instead of telling a new designer "read all 70 pages of the Word file," we could say "memorize the 4.5 pages of L0, and fill in the L2 template for your quest."

One honest addition: the effect did not arrive all at once. In the first quarter we attached only one piece of automation — the auto comment — and everything else was manual. We added interface automation one piece per quarter. Attaching the automation took longer than splitting the Layers.


5.1.11 The Deeper Reason — The Road to Procedural Generation

Everything so far is the surface reason: unify the language of collaboration across disciplines. But there is one more reason, a deeper one. Layer decomposition is the precondition that makes procedural generation possible. The general thesis — each of the five layers maps onto one role in procedural generation (L0 anchor → L1 rulebook → L2 prose → L3 numbers → L4 gate), and when they are mixed into one mass the generator collapses because it cannot decide where to read from and where to write to — was covered in full in §6.6. Here we look only at how that precondition actually operates on the five narrative layers.

On a monolithic 70-page document, a generation algorithm cannot decide where to start reading or where to write. The five narrative layers become, as they are, the five stages of a generation pipeline — L0 the anchor, L1 the input rules, L2 the place where the prose accumulates, L3 the simulation input, L4 the verification gate.

The worked transcript above was already a miniature of these five stages. The L0 pillars and L1 tone were injected as context (the anchor), prose was generated into L2, an L3 row followed, and the tone self-check stood in for the verification gate. Take what a human ran one quest at a time, move it onto the same structure with a generator mass-producing side quests, and you have procedural generation.

Taken further, the flow runs like this — accumulated player actions → state changes in world Behavior Tree (BT) nodes (above the Squad level) → NPC stat changes (reputation, interests, priorities) → NPC tags plus stats become the spawn conditions → matching quests manifest from the quest cloud. Instead of writing every quest in advance, you leave them tagged, "floating in the air like clouds"; when player actions shift an NPC's stats, the quests whose spawn conditions match those stats descend and appear. This progressive model is covered in detail in 5.3 (flow diagram included). The point to stress here is that this entire flow works only on top of the five-layer decomposition.

Conversely, a team whose Layers are mixed together cannot get to procedural generation. The moment they try, consistency accidents bring it down. Just imagine the Kim accident — adoptive father and traitor at once — mass-produced automatically by a generator.

That does not mean you need the full five-drawer chest from day one. In the first quarter, separating just the one L0 line and a single L1 rulebook is enough. Separate gradually; keep the interfaces narrow.

One last note, on timing. Layer decomposition itself is a separation that dates back to the era of deterministic PCG (procedural content generation). What is new is that an LLM now handles natural-language prose, personas, and narrative branching on top of that separation — the AI writing dialogue in Kim's voice in the transcript above was impossible with the rule tables of five years ago. This timing argument is developed further in 5.3.

The next chapter (5.2) shows how lore_consistency_rule automatically verifies worldbuilding→character→quest consistency on top of these five layers — that is, the checker that structurally prevents the Kim accident.


Try It Yourself — Your First Layer Decomposition

Assume you are starting with an existing monolithic worldbuilding document.

setup. Create five empty folders under a NarrativeDocs folder, from Layer0_Vision to Layer4_Build_QA. Leave the existing monolithic document where it is; pick out only the "one line that never changes" and move it into Layer0_Vision/narrative_pillar.md. Put status: locked in the frontmatter. Keep this one file under 4.5 pages.

prompt. When you create new content, give the AI context in this order — the pillars in full, the forbidden words, the relevant bible excerpt, then your one-line idea — and tell it to write the L2 quest file first, fill in the related: block (requires/derives_from/affects), add no lore that is not in the bible, and stop and ask when in doubt.

[narrative_pillar.md in full]
[tone_manifesto.md forbidden words]
[relevant character_bible file excerpt]

Idea: "<one-line idea>"

Do not violate this context, and write the L2 quest file first. Fill in the
related block (requires/derives_from/affects), add no new lore that is not
in the bible, and stop and ask if in doubt.

verify. Check two things in the output. (1) Actually open the files listed in requires: and confirm the AI did not invent facts that are not there. (2) Search the prose for forbidden words (automatic if you have tone_lint.py, by eye if not). If either check fails, reject, state exactly what was wrong, and re-request. Rejecting the "dead wife" in the transcript above was check (1).


5.1.12 Solo Scale-Down

If you are an indie developer working alone, five layers can look like too much. Two drawers will do.

Keep one vision.md (L0 and L1 merged), one content/ folder (L2), and one spreadsheet (L3). Instead of a separate QA layer, substitute one habit: a single requires: line in each content file's frontmatter. Each time you write a new quest, reopen only the files listed in requires: and check against them — you prevent the Kim accident without digging through 70 pages. What matters is not the number of layers but the habit of keeping "the one line that never changes" separate and showing it to the AI first, every time. That one line is your context anchor, and whether you are solo or a mid-size team, that is where you start.


Key Takeaways

5.2 Worldview → Character → Quest Consistency Verification

Just before beta, a bug report came up from QA. The title: "The king is speaking informally." The body was short. "In the 3.4 intro cutscene, K_001 (the King) says '야, 잠깐만' ('Hey, hold on') to the player. This character uses '그대' — the archaic, courtly second-person address — everywhere from 1.1 through 3.3."

I asked the writer, and the answer surprised me. "I never wrote that line." Tracing it back, an outsourced writer had hastily filled in one line of a cutscene branch. Our character bible had a voice_profile, but that outsourced writer had never seen the document. The rule lived inside the document; the line came in from outside it.

This is the essence of a consistency incident. It happens not because there is no rule, but because the rule fails to follow the text all the way down. And if that one line sits in a cutscene, it gets scarier. Cutscenes usually get voice recording. While it is still text, one fix ends it; if it is discovered after recording, the irreversible costs follow — re-booking the voice actors, re-recording, re-mixing. The real goal of consistency verification is to catch it before recording.

This chapter covers a workflow that catches that incident with a rulebook and checkers instead of human eyes. We look at the actual artifacts, as-is: how the lore_consistency_rule rulebook becomes the checker's input, how voice_lint pulls tone drift as suspect candidates, and why the final judgment must remain a human seat to the very end.


5.2.1 Where Consistency Incidents Leak From

Collect user reviews of shipped RPGs and MMORPGs and the narrative consistency incidents converge into a few patterns. The types look different, but the cause is almost always the same.

The five look like different incidents, but trace them and they leak from the same spot. Layer 0 (world premise) or Layer 1 (rules) changed, and that change never propagated to Layer 2 (prose) and Layer 3 (data sheets). The rule got updated; the prose stayed stopped on top of the old rule.

Trying to stop this with manual review is asking too much. One chapter tangles together 50 NPCs, 2,000 lines of dialogue, and 30 quests; when one rule line changes, no human can trace 100% of where its effects spread. The missed line doesn't get caught at the review stage — it gets caught in the user reviews after launch.

That said, automated checking doesn't guarantee 100% either. The point is the division of labor. Automated checks pull suspect candidates fast; humans make the judgment. The goal of automation is to cut human review time, not to remove the human. Blur this premise and every failure covered later in this chapter follows.


5.2.2 lore_consistency_rule — The Rulebook You Feed to the Checker

One of Project A's L1 documents is lore_consistency_rule.md. The document is a guide humans read and, at the same time, input the checker parses. The atoms and affects keys in the frontmatter bind those two roles into one body.

---
title: Lore consistency rules
layer: L1
atoms:
  - lore_check_world_rule
  - lore_check_character_voice
  - lore_check_timeline
  - lore_check_faction_relation
related:
  derives_from: [world_premise, narrative_pillar]
  affects: [main_quest/*, character_bible/*, dialogue_id_table]
---

## 1. World rules
- Magic starts banned → any use must specify (when, by whom, justification)
- The gods are silent → no depictions of direct replies (dreams and visions allowed)

## 2. Character voice rules
- Referencing each character's voice_profile is mandatory
- New dialogue must follow the 5 voice_profile items (vocabulary, sentence length, honorifics, emotional expression, forbidden terms)

## 3. Timeline rules
- Every NPC gets a status_timeline (alive / injured / dead / missing / relocated)
- status_timeline checked automatically at dialogue and appearance points

## 4. Faction relation rules
- Record when the faction_relation_matrix changes
- Dialogue written after a change reflects the new relations

That one affects line defines the checker's scan scope. When world_premise changes, the checker re-sweeps all of main_quest/*, character_bible/*, and dialogue_id_table. The work a human used to do in their head — "how far does this change reach?" — is taken over by the dependency graph written in the rulebook.

voice_profile is a separate L2 asset this rulebook references. A single character's profile is entered as numbers and enumerated values, so the checker can use it as a baseline for comparison.

# character_bible/K_001_voice_profile.yaml
character_id: K_001
display_name: 국왕                       # "the King"
voice_profile:
  vocabulary_register: 고풍_격식        # vocabulary register: archaic_formal
  avg_sentence_len: 18                   # average sentence length (in characters)
  honorific: "그대"                      # fixed second-person address (archaic "thou")
  emotion_expression: 절제               # emotional exposure: restrained
  forbidden_terms: ["야", "잠깐만", "ㅋ"] # forbidden terms: casual "hey" / "wait a sec" / "lol"

Only with this yaml does "the king is speaking informally" stop being human intuition and become items a machine can compare. If honorific is fixed at "그대" and the line contains "야", that's not an opinion — it's a rule-violation candidate.


5.2.3 The Consistency Verification Flow

The moment a change occurs, the checker fires. The flow looks like this.

flowchart TD
    A[Change occurs: L0 premise / L1 rule / L2 prose] --> B{Change classifier}
    B -->|which rule's affects does it hit| C[Invoke the matching checker]
    C --> D[Scan L2 prose + L3 sheets in affects scope]
    D --> E[List of rule violation / suspect candidates]
    E --> F[Auto-attach comments to the change request]
    F --> G{Human judgment}
    G -->|genuine violation| H[Fix the prose → re-check]
    G -->|intended change| I[Update voice_profile / rulebook]
    G -->|rule too sensitive| J[Adjust the rule itself]
    H --> K{Still at the text stage}
    I --> K
    J --> K
    K -->|yes: reversible| L[Review can close]
    K -->|no: after recording| M[Irreversible — re-recording cost]
    L -.cutoff line.-> M

    classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545;
    classDef human fill:#fde68a,stroke:#b45309,color:#000;
    classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b;
    classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d;
    classDef fail fill:#fee2e2,stroke:#dc2626,color:#7f1d1d;
    class B,C,D,F code;
    class G human;
    class A,E data;
    class L pass;
    class M fail;

The last branch is the hidden spine of this chapter. Every consistency judgment must end at the text stage — that is, the reversible stage. Once review slips past recording and casting, fixes become irreversible. That's why checkers like voice_lint and timeline_lint need to run not fast but early. A cutscene line has to pass through at least once before it enters the recording queue.

There are four checkers, each mapping one-to-one to a section of the rulebook.

None of the four checkers is 100% accurate. That's why the output is named "suspect candidates," not "violations."


5.2.4 Worked Transcript: One Full Pass of voice_lint

An abstract "there is a checker" doesn't give you a feel for it. Let's actually run it once. The input reproduces the "king speaking informally" incident from the opening.

setup — Pull the two lines under inspection from dialogue_id_table.

dialogue_id_204  speaker=K_001  text="야, 잠깐만요. 그쪽이 먼저 말해 봐."
dialogue_id_217  speaker=K_007  text="...젠장, 또 실패야. 다시 처음부터."

(Line 204 has the king saying, in a casual register, "Hey, hold on. You go first." Line 217 has K_007 muttering "...Damn it, failed again. Back to the start.")

prompt — The checker itself runs on rules: regular expressions, length comparison, forbidden-term matching. The AI is attached only at one seat — explaining why something is suspect, in natural language a writer will actually read. The full LLM prompt for the inspection step follows.

voice_lint가 뽑은 의심 후보들이야. 판정은 하지 말고, 각 후보가
어떤 voice_profile 항목을 위반한 것 같은지랑 근거만 1~2문장으로 설명해줘.
"위반이다/아니다" 단정하지 말고 "~로 의심됨"까지만.

[voice_profile]
K_001: honorific="그대", avg_sentence_len=18, forbidden=["야","잠깐만"]
K_007: forbidden=["젠장"]

[의심 후보]
- dialogue_id_204 (K_001): "야, 잠깐만요. 그쪽이 먼저 말해 봐."
- dialogue_id_217 (K_007): "...젠장, 또 실패야. 다시 처음부터."

(The prompt, in brief: "These are the suspect candidates voice_lint pulled. Do not judge. For each candidate, explain in one or two sentences which voice_profile item it seems to violate and why. Do not declare 'violation / not a violation' — go no further than 'suspected.'" It then lists the two voice_profiles and the two candidate lines.)

Claude's raw output (reproduced as-is, with no summarizing or cleanup)

[dialogue_id_204 / K_001]
이 대사는 honorific 항목과 충돌하는 것으로 의심됩니다. voice_profile은
2인칭 존칭을 "그대"로 고정하고 있는데, 본 대사는 "그쪽"을 쓰고 존칭 자체가
드러나지 않습니다. 또한 forbidden_terms의 "야"와 "잠깐만"이 모두 등장하여
어휘 격(고풍_격식)에서 벗어난 것으로 의심됩니다. 평균 문장 길이 18자 기준과는
직접 어긋나 보이지 않으나, 어조 자체가 캐주얼해진 점이 위반 신호로 의심됩니다.

[dialogue_id_217 / K_007]
forbidden_terms의 "젠장"이 그대로 사용되어 금기 표현 항목 위반으로 의심됩니다.
다만 K_007(학자)의 emotion_expression 항목이 제시되지 않아, 이 표현이
의도된 감정 고조인지 단순 일탈인지는 이 정보만으로 판단하기 어렵습니다.
판정은 작가 검토가 필요합니다.

(In brief: for 204, the output suspects a conflict with the honorific item — the profile fixes the second-person address to "그대" while the line uses a casual "그쪽" — and notes that both forbidden terms "야" and "잠깐만" appear, drifting from the archaic-formal vocabulary register; it concedes the average sentence length of 18 does not directly deviate. For 217, it flags the forbidden term "젠장" but says that with no emotion_expression item provided for K_007, a scholar, it cannot tell an intended emotional peak from a simple lapse, and that the verdict needs the writer's review.)

This output carries two kinds of honesty we deliberately left in. First, on 204 the AI said the average sentence length "does not directly deviate" — it did not insist that a non-violating item was a violation. Second, on 217 it said "the emotion_expression item was not provided, so this is hard to judge" — it admitted missing information and handed the judgment to a human. If the AI had pushed every suspicion through as "confirmed violation," that would be the more dangerous checker.

verify — The writer receives these comments directly on the change request. The writer makes the judgment.

The same checker pulled both candidates, yet the endings are opposite. One fixes the prose; the other fixes the rule. That a machine cannot make this fork automatically is the heart of the next section.


5.2.5 Why Judgment Is the Human's Seat

There are three reasons the checker stops at suspicion and hands over the judgment.

First, intended violations exist. In chapters where a character breaks down or changes, the voice drifts on purpose. 217 above is exactly that. An auto-reject checker blocks the writer's dramatic intent.

Second, the rules themselves evolve. If the same kind of suspicion keeps getting judged an "intended change," that's a signal the rule isn't keeping up with reality. Check results don't just drive fixes to the prose; they drive fixes to the rulebook too.

Third, new characters and factions need a learning period. A new NPC whose voice_profile has only two or three items filled in will naturally surface lots of suspicions. Turn on auto-reject during this period and writers start seeing the checker as the enemy.

The checker survives only when the boundary between automated checking and human judgment is sharp. Make it auto-reject and within a month the writers will say, "let's turn this off." It's like installing an oversensitive automatic sensor on the office door: the door slams shut every time someone walks through, until eventually somebody rips the sensor out. A checker should not be a device that closes the door — it should be a device that reports "someone passed through here."

One caveat to add. The principle that review must close at the text stage (the cutoff line in the flowchart above) applies to human judgment just the same. The writer's "intended violation" verdict also has to land before recording. A reversal after recording is no longer a checker problem — it turns into a production-cost problem (the full map of the reversible/irreversible boundary is in 5.4.5).


5.2.6 Measurement — Six Months Before and After

On Project A we rolled out the four checkers in stages and measured six months. The figures below are based on actual logs, expressed as direction and ratio rather than absolute values (internal measurement, not an author's estimate).

The last item is the most interesting. With checkers in place, changing rules often is safe. Change one rule line and its impact becomes visible automatically, so the fear of change shrinks and the rules evolve faster. The real effect of consistency tooling is less "we reduced incidents" and closer to "we can now change rules without fear."

One caution: the numbers above are from the point where all four checkers were running. What matters more is that in the early phase, voice_lint alone produced visible results. You don't need to switch on all four from day one.


5.2.7 Where to Put the AI

For the checker core itself, rule-based is the efficient choice. Trust builds only when the same input yields the same result, and an LLM is non-deterministic — the wrong fit for that seat. The AI goes into four other seats.

Rules are fast and deterministic; LLMs are strong at explanation and generation. Mix the two roles and both break. Hand the checking to an LLM and the same line passes yesterday and fails today; hand the explanation to a regex and all you get is machine-speak like "honorific item violation."


5.2.8 Adoption Order and Common Failures

Build all four checkers from the start and the burden arrives before the payoff. The recommended order starts with the cheapest, highest-impact pieces.

  1. Standardize the five voice_profile items (about 1 month) — settle the character_bible format first. This comes before any checker
  2. Minimal voice_lint (about 1 week) — forbidden-term matching only. Blocking a single word cuts post-launch social media incidents by 1–2 per quarter
  3. timeline_lint (1–2 weeks) — death-flag checks. Just catching dead NPCs reappearing is a tangible win
  4. world_rule_lint + faction_lint (1–2 months) — the remaining two
  5. LLM assistance (1–2 more months) — integrate explanation and draft generation

I want to stress that step 2 (voice_lint) alone delivers a lot. The "king's informal speech" incident at the opening was exactly the kind this single step catches.

The failures that recur during adoption are almost fixed, too.

The last item costs more than all the items above it. The other failures lose you time; this one loses you the voice actors' schedule.


The next chapter (5.3) covers writing narrative prose with AI assistance, rather than checking it. We look at how to inject L0 tone and L1 rules as context, so the AI produces answers from our world instead of generic ones.


Key Takeaways

Next Chapter Preview


Try It Yourself

setup — Pick one character from your character_bible and fully fill out the five voice_profile items (vocabulary register, average sentence length, honorific address, emotional expression, forbidden terms) in yaml. Pull 10 of that character's existing lines from dialogue_id_table and gather them in one file.

prompt — Use the inspection-assist prompt from the worked transcript above as-is. The heart of it is two constraints: "do not judge" and "go no further than 'suspected.'" Paste the voice_profile yaml and the 10 lines into the input.

verify — Judge the output's suspect candidates yourself, line by line. If it's a genuine violation, fix the prose; if it's an intended change, add an exception flag to the voice_profile. Also check whether the AI insisted that non-violating items were violations, and whether it admitted missing information. If the AI declares every item a violation, strengthen the "do not judge" constraint in the prompt.

Solo Scale-Down

If you're a solo developer with no four checkers and no rulebook, one prompt can get you the same effect without a checker core. Maintain a voice_profile yaml per character by hand, and every time you write new dialogue, paste that character's yaml plus the new lines into the assist prompt above to collect "suspect candidates." There's no automation, but the core structure — humans judge, AI explains — survives intact. There's only one line to hold: run this review once before anything goes to recording or voice synthesis. The principle of never crossing the reversible stage doesn't depend on team size.

5.3 AI-Assisted Narrative Writing

It was the day I was drafting the first lines for a new side NPC. Into an empty chat window I typed, "Give me five lines for a village blacksmith NPC." Five seconds later the screen showed, "Hero, entrust thy weapon to me." Not a sentence that merely felt familiar — a sentence I could practically name the source of. The same prompt would have returned the same answer for a different game on a different team. What I realized in that moment was not that the model was weak. It was that I had told the model nothing about our game.

AI writes generic fantasy sentences well. What it cannot write is our world's sentences. The difference comes down to one thing: context injection. Send the L0 tone and the L1 rules along with every request, and the line the AI produces shifts from "a sentence I've seen somewhere" to "a sentence from this game." This chapter covers the practice of running that injection as four layers, and at the end it marks the progressive application that lifts the same principle to world-simulation scale — the world BT (Behavior Tree) plus the quest cloud — as the front line of R&D.


5.3.1 What Happens When the Context Is Empty

In narrative work, AI assistance is the area adopted fastest and the area that loses trust fastest. The failure pattern is almost always the same.

Throw in "five quest-opening lines" and back come five generic fantasy variants starting with "Hero, our village..." Ask it to "polish this character's dialogue" and the voice flattens out — every NPC converges on a similar register. Ask for "a chapter 1 synopsis" and you get the average of every RPG synopsis the model has ever seen.

The problem is not the model; it is that the context is empty. A model outputs the average of its training data. If you don't want the average, you have to give it cues that pull it away from the average. The subject of this chapter is how to build those cues and how to inject them.

On the MMORPG project I run (Project A from here on), narrative AI assistance stacks four layers of context in order. The structure from 5.1 — NarrativeDocs decomposed into Layers 0–4 — is reused here, as is, as the unit of injection.

Layer A · System prompt Writer persona · prohibitions (rarely changes, defined once) Layer B · L0 vision world_premise · narrative_pillar · tone_manifesto (≈7,000 tok, cached) Layer C · L1 rules (selective injection) Only the _summary sections of job-relevant rules (cached) Layer D · L2 adjacent text Same character's previous lines · same chapter synopsis (verbatim, changes every time) Task instruction: "3 line options for K_007 at this point"

I don't inject all four layers every time; I pull only the layers the job needs. For a draft of one character's next line, A + B (tone only) + D (that character's 10 most recent lines) is enough. For a new side-quest synopsis, C (the quest structure rules) gets added. For four options of a branch outcome, C (the branching rules) + D (the full text leading up to the branch) gets heavy. It's like reaching into the file tray on my desk and picking out — sized to the job — a persona sheet, a one-line world premise, a page of the rulebook, and a bundle of adjacent text.


5.3.2 One Worked Transcript — K_007's First Emotional Line

Instead of explaining in the abstract, I follow one real request from start to finish. The job: draft three line options for the scene where one recurring character — a scholar-type NPC, internal ID K_007 — has to show emotion for the first time. Starting with the full prompt.

The prompt I sent (Layer A + B (tone) + D + task instruction + output format):

[System]
You are the narrative writer for Project A. No stock RPG phrases like "hero"
or "the chosen one". Follow the tone of the previous lines below. Don't invent
lore you don't know — say "can't judge from this context". You only draft; I decide.

[L0 tone — tone_manifesto excerpt]
The register of this world is restraint. Characters do not speak emotion easily.
Emotion shows as cracks, not outbursts. Ellipses and short sentences.

[Reference — K_007's previous 5 lines (chronological, verbatim)]
1. "...That's the first I'm hearing of it."
2. "Is there a record? If not, we should be skeptical."
3. "Let's not move until we're certain."
4. "That is all I have to say."
5. "...I'm sorry. Was I too cold?"

[Situation]
A fellow NPC has returned wounded. The scholar must show emotion for the first time.

Exactly 3 options, each within 25 characters. End each option with a
(reaction type: anger/sadness/denial) label, and skip explanations and preamble.

(In English, the prompt reads: the system block says — you are the narrative writer for Project A; no stock RPG phrases like "hero" or "the chosen one"; follow the tone of the previous lines below; don't invent lore you don't know — say "can't judge from this context"; you only draft, I decide. The L0 tone excerpt: the register of this world is restraint; characters do not speak emotion easily; emotion shows as cracks, not outbursts; ellipses and short sentences. Then the scholar's five previous lines, verbatim and in chronological order — 1 "...That's the first I'm hearing of it." / 2 "Is there a record? If not, we should be skeptical." / 3 "Let's not move until we're certain." / 4 "That is all I have to say." / 5 "...I'm sorry. Was I too cold?" — the situation: a fellow NPC has returned wounded, and the scholar must show emotion for the first time; and the format: exactly three options, each within 25 characters, each ending with a label (reaction type: anger/sadness/denial), no explanations or preamble.)

Claude's raw output (verbatim — no summarizing, no retouching):

1. "...Why did you carry that alone? Like a fool." (reaction type: sadness)
2. "I should have checked the records. I— I missed it." (reaction type: denial)
3. "It's okay. It's all okay. So stop—" (reaction type: denial)

The point is what happens next — we don't stop here. The writer reads the three options and verifies them. Option 1 — "...왜 그걸 혼자 떠안았어요. 바보같이." ("...Why did you carry that alone? Like a fool." — labeled sadness) — gets the tone right. The crack that previous line 5 opened with "...죄송합니다. 제가 너무 차가웠나요." ("...I'm sorry. Was I too cold?") is carried forward by a single word, "바보같이" ("like a fool"). A candidate for adoption. Option 2 ("기록을 봤어야 했는데. 제가, 제가 놓쳤어요." — "I should have checked the records. I— I missed it.") stammers, and the stammer sits slightly off from the scholar's restraint. On hold. Option 3 ("괜찮아요. 다 괜찮아요. 그러니까 그만—" — "It's okay. It's all okay. So stop—") has the wrong label: repeating "it's okay" is not denial but something closer to avoidance, and our voice_profile has no pattern of this scholar processing emotion through avoidance. Rejected.

The re-request (sent back with the rejection reason stated):

Option 3 rejected. The "it's okay" repetition isn't this scholar's pattern. This character
processes emotion not as avoidance but as 'self-blame over a failed fact-check'. Option 1 is adopted.
In place of options 2 and 3, give me two new options with a different grain from option 1. No stammering, within 25 characters.

(In English: option 3 rejected — the "it's okay" repetition isn't this scholar's pattern; this character processes emotion not as avoidance but as self-blame over a failed fact-check. Option 1 is adopted. In place of options 2 and 3, give me two new options with a different grain from option 1. No stammering, within 25 characters.)

Claude's second output:

2. "Records don't lie. I just didn't look." (reaction type: self-blame)
3. "...Sit down. The wound first. Words later." (reaction type: sadness)

This time both passed. Option 2 — "기록은 거짓말 안 해요. 내가 안 본 거예요." ("Records don't lie. I just didn't look." — labeled self-blame) — reuses the scholar's core vocabulary, "기록" ("records"; see previous line 2, "Is there a record?"), as the vehicle for self-blame. Option 3 — "...앉아요. 상처부터. 말은 나중에." ("...Sit down. The wound first. Words later.") — shows, in restrained imperatives, the scholar's pattern of pressing emotion down into action. The final picks: option 1 plus the new options 2 and 3. These three lines pass the automated voice_lint review from 5.2, get merged into the L2 text, and receive a dialogue_id from L3.

This one transcript holds everything in this chapter. Tone injection (L0) saved option 1; the verbatim adjacent text (L2) let the model pick the scholar's word "기록" back up at the re-request; the enforced output format blocked the chatter; and the writer's rejection gate caught option 3's wrong label. The AI made not a single final decision.


5.3.3 Layer A — The System Prompt: One Line Decides Everything

This is the persona definition laid down on top of everything else. Define it once and almost never change it. The system block in the transcript above is the actual artifact. Of its five lines, the last one ("you draft, the writer decides") matters most. Leave it out and the AI confidently produces sentences that pose as "final," and the writer ends up grading instead of reviewing. The third line ("don't invent lore you don't know — answer that it can't be judged from the context") matters second-most. Without it, the model fills the blanks with plausible lies. In narrative, a plausible lie comes back a few days later as a lore conflict.


5.3.4 Layer B — The L0 Vision and Where the Cache Goes

L0 is small (about 4.5 pages by the count in 5.1). Injecting all of it nearly every time is feasible. Estimated against the Korean text, world_premise.md runs about 2,500 tokens, narrative_pillar.md about 1,500, and tone_manifesto.md about 3,000 — roughly 7,000 tokens combined. (These figures are the author's estimates, unverified. They shift with the tokenizer and with document revisions.)

Sending 7,000 tokens fresh with every request adds up. So I turn on prompt caching. Both Anthropic and OpenAI support it, and on a cache hit the input-token cost drops sharply. The key is to separate, inside the message, what changes from what doesn't.

messages = [
    {"role": "system", "content": SYSTEM_PROMPT},
    {"role": "user", "content": [
        {"type": "text", "text": L0_FULL,      "cache_control": {"type": "ephemeral"}},
        {"type": "text", "text": L1_SELECTED,  "cache_control": {"type": "ephemeral"}},
        {"type": "text", "text": L2_ADJACENT},   # changes every time — not cached
        {"type": "text", "text": TASK_INSTRUCTION},  # changes every time
    ]},
]

L0 and L1, tagged with cache_control, are cache targets; the L2 adjacent text and the task instruction change every time, so they are not cached. Always grouping the cached blocks at the front of the message is what decides the hit rate. If a changing block slips in front, every cache behind it is invalidated. Getting this order wrong is the most common reason caching is on and the bill doesn't go down.

The details on cache hit rates and cost-saving figures are covered in the Part 22 chapter on cost. Here, the one principle to remember is: push what changes to the back.


5.3.5 Layer C — Don't Inject the L1 Rules Whole

The L1 rulebook is big. Put all of it in and the context blows up — and worse, the model loses the point. Pick only the rules relevant to the job, and of those, only the _summary sections.

For main-quest branch outcomes, pick dialogue_branching_rule and faction_relation_matrix. For new NPC dialogue, that NPC's voice_profile and tone_manifesto. For a new lore-dictionary entry, lore_consistency_rule and world_premise. For a side-quest skeleton, quest_template and reputation_model. Selection is done either by hand or by automatic extraction along the wikilink graph (Part 7); when extracting automatically, favor recall over precision. The damage from one missing rule is far greater than the damage from one extra rule.

Instead of injecting the full rulebook text, keep a _summary section at the head of each rulebook file and inject only that.

---
title: Branching rules
layer: L1
---

## _summary
- Branches occur only at chapter ends
- Branches have 2~3 options. 4 or more prohibited
- A branch choice affects reputation by +/-1; an ending branch by +/-3
- Every branch outcome must show its result within 24 hours
- Branches cannot be undone (expose UI recommending a separate save)

## 1. Rules for when branches occur
(detailed explanation, for operators' reference — not injected into the LLM)
...

Five lines of _summary do more for LLM output quality than fifty lines of body text. Models follow short, assertive rules better. Long explanations scatter the model's attention, and scattered attention comes back as rule violations.


5.3.6 Layer D — Never Summarize the Adjacent Text

Previous lines, adjacent quests, the synopsis of the same chapter — this is the most volatile context. For a character's new dialogue, inject that character's previous 10 lines in chronological order; for a mid-chapter quest, the chapter synopsis plus one-line summaries of the other quests in that chapter; for a branch-outcome ending, the full text leading up to the branch plus the choice text. Inject too much and the LLM outputs the average; inject too little and you get generalized output. The workable zone is somewhere between 1,500 and 3,000 tokens (author's observation, unverified).

One core rule: adjacent text goes in verbatim — never processed, never summarized. In the transcript above, the scholar's five previous lines went in untouched, and that is exactly why the model, at the re-request stage, could pick out the precise word "기록" ("records") and reuse it as the vehicle for self-blame. Had those five lines been summarized as "the scholar is cautious and cold," every one of the writer's fine-grained choices would have vanished, and the model would have drifted back to the average. Summarizing doesn't reduce information; it erases decisions the writer has already made.


5.3.7 The Writer Review Workflow — Using the Discard Rate as a Metric

AI output is always a draft. Review passes through a fixed gate.

flowchart TD
    A[N AI outputs] --> B[Pass 1: voice_lint, automated
5.2] B --> C[Pass 2: one writer selects and edits
about 15 min] C --> D{Any option adopted?} D -->|yes| E[Pass 3: narrative lead sign-off
when needed] D -->|0 adopted = all discarded| F[Re-request or write by hand] E --> G[Merged into L2 text
+ L3 dialogue_id issued] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; classDef fail fill:#fee2e2,stroke:#dc2626,color:#7f1d1d; class B code; class A ai; class C,D,E human; class G pass; class F fail;

It is also normal for the writer to pick zero out of N — to discard everything. As option 3 was rejected in the transcript above, a rejection is not a failure; it is proof the gate worked. So I measure the discard rate per writer and per character, and use it as the metric for context-injection quality.

A discard rate of 0–20% means stable operation with sufficient context — leave it alone. 20–50% is the normal operating range — just monitor. If it climbs to 50–80%, recheck whether an L1 rule was left out of the selection. Past 80%, the problem is not any individual rule but the system prompt and persona themselves being off — rewrite Layer A. The discard rate is tallied per writer once a week and shared in the retrospective.

That said, the discard rate is not an absolute metric. A character changing fast (say, at a turning point like K_007's first show of emotion in the transcript above) can run a high discard rate and still be healthy. The number is the start of a conversation, not a verdict.


5.3.8 Security — How to Stop Context Leaks

L0 and L1 are the game's core IP. If sending them as is to an external LLM API feels uncomfortable, the options diverge. Using an external API as is, under a no-training agreement, is fastest but needs legal review. Swapping company names and proper nouns for placeholders before sending adds processing cost and damages naturalness. Self-hosting an open model keeps the data safe but carries a heavy quality and operations burden. A hybrid — keep L0 in-house and send only drafts outside — is complex to operate.

My Project A uses the first option (external API plus a no-training agreement). We tried the second and dropped it: placeholder substitution flattened the text into the shape of "the ○○ scholar of the ○○ kingdom spoke about ○○" and wrecked output quality. Anonymization killing quality is a trade-off that recurs throughout this book (see the anonymization chapter in Part 1). In narrative the damage is especially severe, because proper nouns are the tone.


5.3.9 Common Failures and Their Fixes

Throw a task instruction with no system prompt and you get the average. Lay down the persona and the prohibitions first. Inject all of L0 every time without caching and the cost bleeds. Group the cache blocks at the front. Stuff the whole L1 rulebook in and the model loses the point. Extract only the _summary sections. Summarize the adjacent text and the writer's choices get erased. Quote the originals verbatim. Skip the output format and a good share of responses open with "Here are three candidates:" — followed eventually by the incident where a writer mistakes that preamble for actual copy. Specify count, length, labels. Use AI output as final and the review gate collapses. Always pass it through the writer gate. Don't measure the discard rate and the tool's health depends on people's impressions. Tally it weekly and share it in the retrospective.


5.3.10 From Conservative to Progressive Application

Everything so far has been the conservative application: the writer injects context with care, and the AI only drafts, line by line. The unit of work is small — "three options for this character's next line," "a synopsis for this quest." It is stable, but it has limits in production scale and dynamic responsiveness.

One thing is worth marking first. Procedural generation, world simulation, and dynamic quests are visions game designers have been sketching on paper for the last 20–30 years. Deterministic, rulebook-driven PCG handled the numeric territory — dungeon rooms, weapon options, spawn distributions — but never reached natural-language text, character personas, narrative branching, or NPC dialogue. A large share of design stayed on paper. The advances in LLMs and image models from 2024 to 2026 pulled that territory into implementable range. The core meaning of AI progress is not model scores; it is that designs long stuck on paper became feasible. That said, a possibility opening up and a system settling into something operable are different problems.

Layer Decomposition Was the Precondition for Procedural Generation

The Layer 0–4 decomposition in 5.1 was not mere tidying; it was the precondition for procedural generation. Which stage of the generation pipeline each of the five layers maps to (L0 anchor → L1 rulebook → L2 text → L3 numbers → L4 gate) was covered in §6.6 and 5.1.11. The point is single: on top of one monolithic document, a generator cannot decide where to start reading or where to write; not knowing where L0 sits, its context blurs; L1 and L2 collide inside one file and the production line collapses. Layer decomposition comes first, and procedural generation runs on top of it. Here, that precondition is applied to the mass-production stage of narrative AI assistance.

The Skeleton of the Progressive Application — The Quest Cloud

Three elements come together. First, NPC Personas are procedurally generated. Main NPCs stay in the writer's hands, but side NPCs are mass-produced by the generator and Squad pipeline from 6.2–6.3. Each Persona carries a voice_profile plus tags (occupation, faction, disposition, role). Second, the world BT sits above Squad. Where a Squad bundles the behavior of one hunting ground's NPC group, the world BT sits above it, takes accumulated player-action metrics, and updates the state of the world as a whole. Which factions the player helped, which regions they frequented, which decisions they made — these shake the world BT's node states, and the shaken nodes propagate into the stats (reputation, interests, priorities) of the NPCs in their sphere of influence. Third, quests follow a cloud model. Every quest except the main quest is procedurally generated and floats free, each carrying tags (who, where, why, when) and manifestation conditions. No quest is pinned to a specific NPC.

Manifestation passes through five stages.

flowchart TD
    P[1. Player actions accumulate
faction support · region visits · decisions] --> W[2. World BT node states change
world state updated above Squad] W --> N[3. NPC stats change
reputation · interests · priorities] N --> M[4. NPC tags + stats == manifestation conditions
matching key generated] M --> Q[5. Matching quest manifests
from the quest cloud] Q -.->|shown after passing the writer gate| PL[Appears to the player] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; class M code; class P,W,N,Q data; class PL pass;

As a metaphor: quests are not books shelved in a library but clouds drifting overhead. When an NPC reaches a certain state, the cloud that fits it descends within reach. In this structure the AI is no longer an assistant writing one line; it handles three things at once — the natural-language text (descriptions, dialogue) of the procedurally generated Personas and quests; the vocabulary shifts when world BT state lands on an NPC (the same voice_profile holds, but the topics move with the world state); and the suspicion classifier at the verification stage (should this quest be allowed to manifest on this NPC?).

The Irreversibility Boundary — Review Must End at the Text Stage

Even in the progressive application, the reversible/irreversible boundary (5.4.5) stays fully alive. If a cloud-manifested quest's dialogue flows into the voice pipeline without passing review, no code rollback can bring it back. So the safety mechanism of the progressive application is placing every review gate (voice_lint, the suspicion classifier, the writer gate) in front of the irreversibility boundary. With automated mass production added on, these gates have to be harder than in the conservative application.

Where to Stop, and Why Not to Stop

This book does not cover the finished form of the progressive application. That is a book of its own, and the infrastructure assumptions differ by company and project. Only two things need remembering. First, a team that can't run the conservative application can't run the progressive one either. If the review gate doesn't turn in the conservative application, the cloud runs away in the progressive one. For the operation to survive, five things must be in place together: procedural generation infrastructure; tooling to define and test world BT nodes; automatic suspicion classification of manifested quests plus the writer gate; simultaneous tracking of manifestation rate, discard rate, and player satisfaction; and automatic recall and replacement of wrongly manifested quests. Drop any one of the five and the system gets scrapped within a quarter — consistency incidents explode past what human review can keep up with.

Second, the progressive application is not a tool for scaling up mass production; it is a tool for scaling up dynamic responsiveness. Aim only at volume and the cloud fills up with generic-RPG average, and players meet a world that "looks more varied but is more empty." That doesn't make this road purely dangerous, either. This territory is the front line of game design R&D. The spot where procedural generation, simulation, and LLMs meet is where a genuinely new game form is most likely to appear. Few companies will adopt it tomorrow, but it is worth keeping in view as one direction for the next 5–10 years.


5.3.11 Try It Yourself: One Chapter End to End

Here is the procedure for reproducing this chapter's conservative application as is.

setup. Merge the three L0 files (world_premise.md, narrative_pillar.md, tone_manifesto.md) into one text and put it in an L0_FULL variable. Write the system prompt exactly as the five lines in the transcript above, and do not drop the last line ("you draft, the writer decides"). Collect the previous 10 lines of the character you're drafting for, verbatim, as plain text.

prompt. Assemble the message in the order [시스템] → [L0 톤] → [참고: 직전 대사 원본] → [상황] → [출력 형식] (system → L0 tone → reference: previous lines verbatim → situation → output format). In the output format, always state "exactly N options / max length per option / labels / nothing else." Put cache_control on L0 and L1 only, and push the changing blocks to the back of the message.

verify. Verify the N options you receive, line by line. Check that the tone continues from the previous lines, that the labels match the actual reactions, and that no pattern your character never uses (avoidance, stammering, and so on) has crept in. For options you reject, state the reason and re-request. Adopting zero is normal; record it in the discard rate. Only options that pass go through voice_lint, get merged into L2, and receive a dialogue_id from L3.

Solo Scale-Down. Without an API, caching, or voice_lint, the core still holds. One ChatGPT or Claude chat window is enough. Paste the character's previous 5–10 lines at the very top every time, always attach the three lines of output format, and for any answer you reject, write down the reason and re-request. Instead of caching, keep the same conversation thread and the earlier context stays in place. For the discard rate, a sheet of paper with a hand tally like "adopted 12 / out of 30 this week" is plenty. The tools are smaller, but the skeleton — four-layer injection and a review gate — works the same for a solo writer.


Key Takeaways

Next Chapter Preview

5.4 Dialogue and Voice Consistency

A recording booth. The director, headphones on, raises a hand to give the cue. The voice actor reads a line for a scholar NPC: "와, 진짜 대박이네요" — "Wow, that's seriously awesome." The director's hand freezes. This is a character who has not used the interjection "wow" once across 5 chapters — a scholar who never utters profanity or modern slang, who trails off at the end of his sentences rather than finishing them. And yet there the line sits, right in the script.

Two costs land at the same time here. First, if the line gets fixed, the actor has to read it again — and the session fee and the studio clock are already running. Second, and scarier, is the case where the director fails to catch the line at all. Once the recording wraps and the audio goes into the build, that scholar talks that way in the game forever. Recorded audio cannot be rolled back with a one-line edit the way text can. You would need the same actor, in the same condition, in the same booth, all over again.

This chapter is about the dotted line in front of that booth. Above the line everything is text, revisable without limit; below the line everything is audio, beyond fixing. Every dialogue review must finish above that line. For that to happen, "this is how this person talks" has to be recorded in a file rather than in someone's head, and every new batch of dialogue has to be checked against that file automatically. I call the file voice_profile, and the checking tool voice_lint.


5.4.1 Where the Voice Starts to Drift

Of all narrative-consistency incidents, character voice drift blows up the most often. Only after launch do the reports arrive: "why did this NPC suddenly change the way they talk?" The cause is nearly the same every time.

When the writer changes, the same NPC becomes a different person. Even when the writer stays, six months later they forget their own tone. Writing new lines without opening the character's previous lines breaks the context. If the voice rules live only in the writer's head and never in a document, they don't transfer to the next person. And then the fifth cause — the one that has grown fastest over the past two years: throw "write some lines for this NPC" at an LLM with no character information, and the AI returns dialogue that is average, inoffensive, and therefore nobody's voice.

That fifth cause is the side effect of adopting AI. Context injection from the previous chapter (5.3) is the prescription, but if the context you inject is thin, there is nothing to inject. That context is the voice_profile. This chapter walks one full cycle: build the file, check new dialogue against it automatically, and close everything out before the recording booth.


5.4.2 voice_profile — A Voice Pinned Down in Five Fields

On Project A, every NPC has a voice_profile with five fields: vocabulary range (word groups used often / word groups never used), sentence length (average and maximum character count), forms of address (first and second person, ratio of honorific — formal — speech, titles), emotional expression (direct, indirect, or suppressed), and taboo expressions (words and phrasings never used). All five fields must carry concrete examples. An abstract description like "a grave scholar" alone gets read differently by everyone. The next writer will imagine that scholar their own way.

Here is the actual profile format for the scholar NPC K_007. This file is, directly, the input to voice_lint.

---
title: K_007 Scholar voice_profile
layer: L1
character_id: K_007
atoms:
  - voice_profile_k_007
related:
  derives_from: [character_bible/k_007.md]
  affects: [dialogue_id_table (all K_007 dialogue)]
---

## 1. Vocabulary
- Frequent: "record", "evidence", "circumstances", "inference", "data", "case"
- Never: "feeling", "gut sense", "fate", "God's will", "the heart's voice"

## 2. Sentence Length
- Average: 18 characters
- Maximum: 35 characters (anything longer splits into two sentences)
- Frequent short cutoffs: "...No." "The records first."

## 3. Forms of Address
- First person: the humble "저"
- Second person: title first ("Captain", "Commander"). Name only after closeness.
- Honorific speech 100% (except flashback scenes)
- Interjections almost never. When present: "...ah."

## 4. Emotional Expression
- Almost no direct expression (anger, joy)
- Expressed through silence and trailing off ("...if that is how it is going to be.")
- Grief: dodged by changing the subject ("...let's talk about something else.")

## 5. Taboo Expressions
- All profanity
- Modern interjections "wow", "whoa", "awesome"
- Two or more Sino-Korean terms of 3+ syllables in a single sentence
- Mysticism vocabulary such as "fate" and "prophecy"

(In English, the profile reads: 1. Vocabulary — frequent: "record," "evidence," "circumstances," "inference," "data," "case"; never: "feeling," "gut sense," "fate," "God's will," "the heart's voice." 2. Sentence length — average 18 characters, maximum 35 (anything longer splits into two sentences); frequent short cutoffs like "...No." "The records first." 3. Forms of address — first person: the humble "저"; second person: title first ("Captain," "Commander"), name only after closeness; honorific speech 100% (except flashbacks); interjections almost never, and when present, only "...ah." 4. Emotional expression — almost no direct expression of anger or joy; shown through silence and trailing off ("...if that is how it is going to be."); grief dodged by changing the subject ("...let's talk about something else."). 5. Taboos — all profanity; the modern interjections 와/헐/대박 ("wow"/"whoa"/"awesome"); two or more long Sino-Korean terms in a single sentence; mysticism vocabulary such as "fate" and "prophecy.")

The heart of this format is that every abstract slot has an example line attached. Not "suppresses emotion" but the actual line "...그런 식이라면." — "...if that is how it is going to be." That is what lets the next writer, the translator, and voice_lint all look at the same standard. Viewed again through the eyes of translation and localization, the five fields split into two kinds. Language-dependent properties (vocabulary range, sentence length, the surface forms of the address system) must be re-specified for every target language; language-independent properties (whether the character suppresses emotion or expresses it directly, what they refuse to say to the very end — attitude) must hold no matter which language the lines move into. Hand this distinction over along with the localization assignment, and you stop a translator from shaking the character's attitude while swapping out surface vocabulary.


5.4.3 The Worked Transcript: Drawing the Profile from the Lines

Writing 50 NPCs × 5 fields = 250 entries from a blank page drifts into abstraction. With not one actual line on the page, "this character is a cold scholar" tells nobody what that coldness is. So the order is flipped: only the 5–7 main NPCs get full profiles from the start; everyone else waits until 20–30 lines of dialogue have accumulated, and the profile is reverse-extracted from those lines.

Below is the actual worked transcript of pulling a profile draft after 25 lines of K_007 dialogue had piled up. Nothing is summarized, and the spots where the AI got things wrong are left in.

[The writer's prompt, in full]

The following are 25 lines of dialogue for character K_007.

1. "Let's start with the records."
2. "...The circumstances don't line up."
3. "It's only an inference. There's no evidence."
4. "I know that case as well."
5. "...No. The data says otherwise."
   (... lines 6~23 quoted in full, no omissions ...)
24. "...If that is how it is going to be, I will step out."
25. "Let's talk about something else."

From these, pull five things: frequently used vocabulary (only what actually appears),
average sentence length, honorific and address patterns, the mode of emotional
expression, and expression ranges never used once. For each item, cite 2~3 numbered
lines as evidence, and do not use subjective impressions like "grave" or "cold".

(The prompt lists K_007's 25 lines in full — "Let's start with the records." / "...The circumstances don't line up." / "It's only an inference. There's no evidence." / "...If that is how it is going to be, I will step out." / "Let's talk about something else." — then asks for five things: frequently used vocabulary (only what actually appears), average sentence length, honorific and address patterns, the mode of emotional expression, and expression ranges that never appear. Each item must cite 2–3 numbered lines as evidence, and subjective impressions like "grave" or "cold" are banned.)

[Claude's raw output — not cleaned up]

1. Frequently used vocabulary
   - "record" (line 1), "circumstances" (line 2), "inference/evidence" (line 3),
     "case" (line 4), "data" (line 5)
   → Concentrated on fact-and-evidence nouns.

2. Average sentence length
   - About 14.6 characters across the 25 lines. Longest 21 characters (line 24).
   → Mostly short declaratives.

3. Honorifics and address
   - First person "저" (lines 3, 24). Polite "-습니다/-요" sentence endings.
   - No second-person address appears in these 25 lines.
     → The address pattern cannot be determined from this sample alone. (Writer confirmation needed)

4. Emotional expression
   - No direct emotion words. Refusal handled as "I will step out" (24),
     grief/avoidance handled as the subject change "Let's talk about something else" (25).
   → Suppressed, indirect type.

5. Ranges never used once
   - Interjections, profanity, and mysticism vocabulary: all unobserved.
   ※ But "unobserved" is no guarantee of "taboo".
      The sample is small at 25 lines; the absence may be coincidental.

(The output, in English: 1. frequent vocabulary concentrates on fact-and-evidence nouns — "record," "circumstances," "inference/evidence," "case," "data" — each cited by line number; 2. average length about 14.6 characters across the 25 lines, longest 21 (line 24) — short declaratives; 3. first person "저" with polite "-습니다/-요" endings, but no second-person address appears in these 25 lines — "the address pattern cannot be determined from this sample alone (writer confirmation needed)"; 4. no direct emotion words — refusal handled as "I will step out" (24), grief/avoidance as the subject change "Let's talk about something else" (25) — suppressed and indirect; 5. interjections, profanity, and mysticism vocabulary all unobserved — with the caveat that "unobserved" is no guarantee of "taboo," since 25 lines is a small sample and the absence may be coincidental.)

Two things deserve a close look here. First, the AI stopped at item 3. It drew its own limit — "cannot be determined from this sample alone." No second-person address appeared in the 25 lines, so it refused to infer one. That is honest output. Had the AI filled in "second person: presumably 'Captain,'" that would be a hallucination, and once inside the profile, contamination. Second, in item 5 the AI attached its own disclaimer. "Unobserved ≠ taboo" is the key warning. The absence of interjections in 25 lines may be pure chance.

[The writer's verification and veto]

The writer adopted items 1, 2, and 4. For item 3, the address pattern, the writer opened the character_bible and filled in "title first ('Captain'), name only after closeness" by hand — a human fills the slot the AI left empty. For item 5, exactly as the AI warned, the writer did not promote "unobserved" straight to "taboo." Instead, checking against the character's design, the writer confirmed only "modern interjections, profanity, mysticism" as taboo and put the rest of the unobserved vocabulary on hold.

[The writer's follow-up request]

I'm confirming "fate" and "prophecy" as taboo. Pull 10 more mysticism words with overlapping meanings.
But since K_007, as a scholar, might quote them in a rebuttal or critique context,
mark that exception case in one line as well.

(The follow-up: "I'm confirming 'fate' and 'prophecy' as taboo. Pull 10 more mysticism words with overlapping meanings. But since K_007, as a scholar, might quote them in a rebuttal or critique context, mark that exception case in one line as well.")

This last follow-up is the important one. Broaden a taboo mechanically and you also block legitimate lines — the scholar deriding superstition with "fate, of all things." So the exception context is defined together with the taboo. The AI widens the candidates; the writer draws the boundary. Only after this full loop does voice_profile_k_007 get confirmed and pinned at L1.

Forcing evidence citations ("cite by line number") and banning subjective adjectives ("no 'grave'") cuts AI hallucination and gives the writer a surface to verify. The profile is not something the AI writes — the AI lays down a draft, and the writer pins it.


5.4.4 voice_lint — Automatic Checks on Every New Line

Once the voice_profile exists as a file, every new batch of dialogue can be checked against it automatically. Of the five possible checks, the two that earn their keep in practice are taboo vocabulary matching (does a word hit the taboo list?) and vocabulary range violation (does it fall into the never-used word group?). Sentence-length deviation, missing honorifics, and frequent-vocabulary ratio throw too many false positives, so they run only as auxiliary checks. Flag every flashback line that bends the average length, and the writer goes numb to warnings.

voice_lint takes a chapter's batch of new lines and produces a report like this.

voice_lint result (ch04 new dialogue: 32 lines, profile=voice_profile_k_007)
─────────────────────────────────────────────
[VIOLATION] dialogue_id_412 — K_007
  Content: "Wow, that's seriously awesome!"
  Reason: taboo words "wow", "awesome" (profile §5)
  → Writer review required

[SUSPECT] dialogue_id_421 — K_007
  Content: "That fate is hard to accept."
  Reason: taboo word area "fate" (profile §5)
        But the 'rebuttal/critique context' exception may apply — writer's call
  → Writer review required

[CLEAN] 30 lines
─────────────────────────────────────────────
Summary: 1 violation / 1 suspect / 30 clean

(The report covers 32 new lines in ch04 against voice_profile_k_007. One violation — dialogue_id_412, "Wow, that's seriously awesome!", taboo words 와 and 대박 (profile §5), writer review required. One suspect — dialogue_id_421, "That fate is hard to accept.", touching the taboo word area "fate" (profile §5) but possibly covered by the rebuttal/critique exception — writer's call, writer review required. 30 lines clean. Summary: 1 violation / 1 suspect / 30 clean.)

Violations are red, suspects are yellow. Both must pass the writer's judgment to get through. And one absolute principle applies here — voice_lint never rejects automatically (an extension of the 5.2 principle). Look at dialogue_id_421 above. "Fate" is taboo, but if the scholar is rebutting a superstition, it can be a legitimate quotation. A tool cannot make that call. An auto-rejecting lint walls off every one of these delicate spots — and robs the writer of the chance to refine the tone on top of it. Lint is a flashlight that marks the suspect spots, not a lock on the door.


5.4.5 Review Gates — The Dotted Line Between Reversible and Irreversible

The spine of this chapter is this single diagram. As one line of dialogue travels from the writer's hands to the actor's mouth, review gates are laid at every stage. And through the middle of that flow runs a thick dotted line.

flowchart TD
    A["Writer drafts lines (L2)"] --> B["voice_lint automatic
violation/suspect report"] B --> C["Writer self-review (15 min)"] C --> D["Narrative lead sample review
10% sample per chapter"] D --> E["L3 dialogue_id issued
+ translation key mapping"] E --> F["Translation/localization review"] F -.->|"━━━ Reversible / irreversible boundary ━━━
Above this line: text, endlessly revisable
Below this line: audio, no revisions"| G G["VA casting"] --> H["Dubbing session (final, irreversible)"] H --> I["Audio into build"] classDef reversible fill:#e8f4ea,stroke:#3a7d44,stroke-width:1px,color:#1b3a22; classDef irreversible fill:#f7e3e3,stroke:#b23b3b,stroke-width:2px,color:#5a1414; class A,B,C,D,E,F reversible; class G,H,I irreversible;

Green marks the reversible stages, red the irreversible ones. Everything above the dotted line (green) is text. If a line bothers you, fix it at the keyboard; the cost is a few minutes of the writer's time. Below the line (red) is audio. The moment the actor reads that line in the booth and the audio enters the build, the line is frozen as an asset. To change it you must rebook the same actor, the same condition, the same studio — and the session fee, the studio, and the director's time all cost exactly what they did the first time. On a tight schedule, an additional session with the same actor often cannot be booked at all.

So a single rule governs the entire workflow — every review gate finishes above the dotted line. Recording is not a review stage. It is the stage that freezes fully reviewed output into an asset. If a "this line sounds off" doubt surfaces below the line, that is not a place to review further — it is a signal that an upstream review was skipped. The flashlight must sweep everything above the line. The booth is not a place that is allowed to be dark; it is a place that must not be.

In the middle of the diagram, the lead's sample review is pegged at a "10% sample per chapter." That ratio is the balance point between review time and accuracy (the author's operating figure — an unverified estimate). Drop below 5% and incidents start leaking; push above 20% and one lead becomes the bottleneck. Because lint has already filtered out violations and suspects, the sample is drawn from the lines that passed lint — human eyes concentrate on the context errors a tool cannot catch (for example, a "fate" quotation that looks legitimate but is in fact the character breaking).


5.4.6 Characters Change — Version Control for the Profile

A character who talks the same way to the very end stalls the plot. A scholar who has lived through a comrade's death and still speaks in exactly the pre-loss tone reads as fake. When the change is intended, the voice_profile versions up along with it.

---
character_id: K_007
voice_profile_versions:
  - v1: ch01~ch05 (early — suppressed emotion, short-sentence scholar)
  - v2: ch06~ch10 (after a comrade's death — emotional expression more frequent)
  - v3: ch11~ (after the awakening — direct speech appears)
---

(The versions: v1, ch01–ch05 — the early suppressed-emotion, short-sentence scholar; v2, ch06–ch10 — emotional expression grows more frequent after a comrade's death; v3, from ch11 — direct speech appears after the awakening.)

Each version is a separate profile file, and voice_lint reads the chapter number of the dialogue under inspection to choose which version to apply. Hold ch07 dialogue up against v1's "suppressed emotion" rule and every perfectly sound change line lights up as a suspect. Change is not a bug; it is design.

Three signals raise a version. When the writer shakes the tone deliberately, they propose a new version and align with the lead. When voice_lint's suspect count keeps climbing for one character, the writer is shifting the tone without noticing — a signal that the version update is due. When a change event (death, awakening, betrayal) is added to the character_bible, a profile-update alert fires. That said, a character who changes every chapter loses coherence, so a realistic count is 2–4 versions per character.

[Directional signpost — reading the space between characters as a voice space (still premature)] Where voice_lint uses rules to protect consistency within one character, a "voice space" — embedding each character's full set of spoken lines as a single point — looks at whether the distance between characters stays wide enough. If the points drift into a cluster, that is a direct measurement of the voice flattening and convergence that §5.3.1 and §5.4.1 pointed to. But do not fix the distance threshold as an absolute number; read it only as a directional signpost that says "clustering is underway," and it is not a verdict gate that replaces voice_lint (the idea sits in the same spot as the dimensional-vector compression of §8.2.7, and the conceptual intuition is laid out as a single map in Appendix M — a signpost, not a prescription).


5.4.7 Consistency Units That Multiply with Localization and VA

Once dialogue is translated into multiple languages and voiced by actors, the units under management multiply. One Korean line splits into English and Southeast Asian languages, and each of those carries its own tone.

The most frequent leak in translation consistency is the same expression being translated differently from chapter to chapter (caught by a translation-memory consistency check). After that come the character's voice_profile not being reflected in the translation (fix: attach a per-character translation guide) and new vocabulary missing from the glossary (fix: a glossary lint). The translation guide is generated automatically from the voice_profile — "this character: 100% formal speech, no interjections, mysticism vocabulary taboo" is stamped onto the head of the translation brief. The translator carrying that scholar into English sees the same boundaries.

VA (voice actor) review is the review at the last text-reversible stage, just before the work touches what lies below the dotted line. Tone consistency (the intensity of anger and grief) is checked by the director and narrative; pronunciation accuracy (proper nouns) by the glossary owner; breathing and breaks (profile instructions like "frequent short cutoffs") by the director. Results are logged pass/reject in voice_review_log.md (L4) and consulted at the next character's casting.

Rejections are settled, wherever possible, before casting and recording. A script error discovered inside the booth collapses that day's session outright and shakes the schedule of the next one. That does not make dragging out review while postponing the recording the answer, either. If reviews keep jamming right in front of the booth, the upstream workflow (writer, lead) is what is running late — not the recording schedule.


5.4.8 Six Months of Measurement and Cost

On Project A, I tracked six months before and after introducing voice_profile + voice_lint. The absolute counts are the author's estimates (unverified); trust only the direction and the ratios.

Item Before After Direction
Voice incidents per chapter (post-launch) 5–8 1–2 about 1/4
Chapters for a new NPC's voice to settle 3 chapters 1 chapter 1/3
NPCs managed per writer about 15 about 40 about 2.5x
Translation consistency incidents (per chapter) 10–15 2–4 about 1/4
Voice review time (per chapter) 3 days 1 day 1/3

The most meaningful row is NPCs managed per writer. "About 2.5x" does not mean writers were cut; it means the same writer can carry more NPC variety per chapter. The world gets more crowded.

Looking at the cost structure, operating cost is far smaller than adoption cost. Writing the voice_profiles took the writer 2 weeks for the 7 mains; the voice_lint tool took 1–2 weeks of development and 1 day a month of maintenance. On the operating side: 15 minutes of writer self-review per chapter, about 2 hours of lead sample review (the 10% sample), and 1–2 days per character for profile updates in change chapters. Operating cost has to stay small for a system to survive. A tool that is heavy to operate gets quietly retired within a quarter.


5.4.9 Common Failures

Pattern Prescription
Profile holds only abstract description ("grave") Force actual example lines in all 5 fields
Aiming to write all 50 NPCs in full from the start 7 mains in full + reverse-extract the rest from accumulated lines
Auto-rejecting voice_lint Violation/suspect flags + writer judgment. Only a human rejects
Broadening taboos mechanically Record the exception context with each taboo ("quotable in rebuttal")
Profile not updated when the character changes Version it (v1, v2, v3); apply by chapter number
Profile not handed to translation Auto-generate the translation guide from the profile
Trying to fix dialogue after recording Recording is irreversible. Reviews end above the dotted line
Compressing review into the recording schedule Solve it by improving the upstream workflow
Keeping the profile in your head File it, no exceptions. Heads are lost when writers change

Try It Yourself — One voice_lint Cycle

This is the minimum procedure for one full pass with a single profile when a new chapter's dialogue comes in.

setup 1. Open the target character's voice_profile_<id>.md. If it doesn't exist, gather 20–30 of the character's existing lines. 2. Prepare the new lines as a plain-text batch in id / 캐릭터 / 내용 (id / character / content) format.

prompt

This is §5 (taboo expressions) of K_007's voice_profile.
[paste the taboo list]

32 new lines of ch04 dialogue.
[paste in id / content format]

Classify each line as [VIOLATION] (directly contains a taboo word) / [SUSPECT] (touches a
taboo area, but an exception context is possible) / [CLEAN]. Table the violations and
suspects with id, content, and reason. I make the call, so don't reject anything automatically.

(The prompt pastes the §5 taboo list from K_007's voice_profile and ch04's 32 new lines, then asks the AI to classify each line as [위반] violation (directly contains a taboo word) / [의심] suspect (touches a taboo area, but an exception context is possible) / [정상] clean — violations and suspects tabled with id, content, and reason — and ends with "I make the call, so don't reject anything automatically.")

verify 1. Look at the violations ([위반]) first. If one is clear-cut, fix the text (it is above the dotted line, so it's free). 2. Judge the suspects ([의심]) by context. A legitimate quotation passes; anything else gets fixed. 3. Send 10% of the passing lines to your lead as a sample to filter context errors once more. 4. Only after every judgment is in, issue dialogue_ids and push to the recording queue. No further checking in front of the booth.

Solo Scale-Down — If you're a solo developer who can't build the tool, write only one field of the voice_profile per character: §5 (taboo expressions). Every time you write new lines, paste that taboo list at the head of your prompt and tell the AI, "flag only the lines that hit this list." No tool, one line of prompt, and you get 80% of what lint gives you. Run this single pass before handing off to recording (or TTS), and the dotted line in front of the booth holds.


Key Takeaways

Next Chapter Preview

Part 6 · Content Design

6.1 Procedural Content Generation and AI — The One Cell Where the Two Axes Cross

Monday morning, design meeting. A single line on the whiteboard: "1,000 side quests by launch." Someone starts tapping a calculator. One writer spending a day per quest: four years. Put five writers on it and it is still close to a year. The air in the room goes heavy. I have been sitting in this room for 24 years, and I know that in front of that number, people always split into the same two camps. One side says "cut the volume"; the other says "stamp them out with tools." And almost always, the decision was both.

Procedural content generation (PCG) is the old answer on the "stamp them out with tools" side. Dungeon room layouts, weapon option combinations, and enemy spawn pools have been automated with rulebooks and probability tables for twenty years. What is new is not PCG itself — it is that large language models (LLMs) and generative models have taken the seats where natural language, images, and narrative go.

But what this book wants to say is not "attach AI to PCG." Everyone does that. The question is where you attach it. Take one piece of content: unless you pin down, in a single cell, which intensity of automation and which layer of structure it meets at, you end up with a tool that has no seat. This chapter shows how to draw that cell as a coordinate, and what it looks like when one piece of content actually makes a full lap around the pipeline on that cell.


6.1.1 Where PCG Stood Still

Traditional PCG is strong at determinism. The same input produces the same output, and you can verify it. That is why dungeon room graphs, weapon option prefixes and suffixes, and enemy spawn distributions settled in early. "Flaming Sword +5" was coming out automatically twenty years ago.

The problem was always the seat right next to it. The rooms got placed, but the names, looks, and short backstories of the NPCs inside them stayed in writers' hands. "Flaming Sword +5" came out, but the one-line hook — "the last sword the king lost" — did not. A quest generator could draw combinations of objectives and rewards, but "why am I doing this quest" was still written by a person.

On large games, this seat was always the bottleneck. The ratio of mass-producible work to work that needed human hands was roughly 4 to 6, and that 6 ate most of the schedule. Even when the production line cranked out the 4 quickly, if the 6 could not keep up, the whole cycle was capped at that speed.

That is exactly where LLMs and image models come in. The mass-producible range extends into the natural-language, narrative, and visual territory that rulebooks could never handle. That does not mean handing the seat over to AI wholesale. AI gives a slightly different answer every time, and when the context is empty, it spits out the generic-RPG average. So you need to design the junction point. And the junction point is defined by two coordinate axes.


6.1.2 The First Axis: Automation Intensity (L0–L3)

The vertical axis is the ratio in which you mix humans, rulebooks, and AI. At the MMORPG studio where I work (hereafter "Project A"), we cut it into four levels.

L0 — fully handcrafted. Every word and every decision comes from human hands. Main quest text, signature character dialogue, branching endings. The seats where consistency and narrative depth tie directly into the game's identity.

L1 — rulebook automation. Traditional PCG's seat. Deterministic algorithms — rulebooks, probability tables, BSP (binary space partitioning) — produce the output, and humans only review it. Dungeon room layout, weapon option combinations, and enemy spawns are the representatives.

L2 — rulebook + AI assist. The rulebook sets the skeleton and AI fills in the details. Side quest synopses, common-NPC names and short backstories, hunting-ground blurbs. Humans are responsible only for the input metadata and the final review gate.

L3 — AI first + human review. AI writes the body and humans step in only to review. Attractive, but non-determinism, hallucination, and the risk of consistency damage all gather here.

The key is L2. It combines L1's stability with L3's production power, and blocks the weaknesses of both with a verification gate — a quality gate where a human or an automated checker verifies the output before it moves on. L3 is the level everyone is tempted to adopt early, but I have watched it get scrapped within a quarter or two, more than once, as the review load explodes. When 70 out of 100 items come back flagged as suspect, it costs more than having a person write all 100 from scratch.


6.1.3 The Second Axis: Layer Structure (L0–L4)

The vertical axis alone will not make a production line run. The content itself has to be decomposed into layers before automation has a seat to take. This is the Layer decomposition covered in Part 5, and in the content domain it is the horizontal axis. The general explanation — that the five layers each correspond to one role in procedural generation (anchor, rulebook, body, numbers, gate) — was covered in §2.3.6; here I plug it straight into the content production line. Layer 0 vision is the tone and worldview anchor (injected on every generation); Layer 1 systems is the generation rulebook (rules, probability tables, tag taxonomy); Layer 2 content is the body where generated output accumulates (side quests, NPC backstories, city blurbs); Layer 3 data is numbers, IDs, and relations (rewards, spawns, curves); and Layer 4 build/QA is the verification gate (lint, consistency checks, writer review).

These two axes tell different stories. The vertical axis says "how much do humans touch this"; the horizontal axis says "which part of the content is this." But they only carry meaning as a product. Only when you pin a piece of content to the intersection of the two axes — to one cell — does "who makes this, which part, and how" actually get decided.


6.1.4 The Two Axes on One Page — The Automation × Layer Matrix

Now I overlay the two axes I have been describing in prose onto a single grid. Horizontal: the content's Layer. Vertical: automation intensity. The label in each cell is the content that actually occupies that cell on Project A. The darker the cell, the closer it sits to the production line's center of gravity.

Automation intensity (vertical) × Layer structure (horizontal) → Layer structure (which part of the content is this) ↑ Automation intensity (how much humans touch it) L0 Vision L1 Systems (rulebook) L2 Content (body) L3 Data L4 Build/QA L0 Handcrafted L1 Rulebook L2 Rulebook+AI L3 AI-first Tone line written by hand Main quests Signature dialogue Dungeon room layout Option/spawn odds tables Reward curve computation Define generation rulebook Side quest skeleton NPC short backstory/blurb ★ Center of gravity lint/consistency checks Patch note drafts Writer review gate

This grid is the heart of this chapter. Judgments that used to be scattered through prose — "main quests are L0," "side quests are L2," "rewards belong to the rulebook" — converge into one coordinate. When a new piece of content comes up in a meeting, one question is enough: "which cell is this?" Once the cell is fixed, its vertical coordinate tells you who touches it, and its horizontal coordinate tells you which part it is.

Read the grid for a while and two things stand out. First, the center of gravity (the dark cell) sits at the L2 row × Layer 2 column. Side quest skeletons and NPC backstories live there. That is the heart of the production line. Second, a piece of content does not live in only one cell. A side quest's body (Layer 2) sits in the L2 cell, but its reward values (Layer 3) drop down to the L1 cell. Even within the same quest, different parts live in different cells. That is why the two axes are kept separate.


6.1.5 A Mini Pipeline Running on a Single Cell

Now watch one piece of content make an actual lap on the center-of-gravity cell — L2 row × Layer 2 column, the side quest skeleton. The flow looks like this.

flowchart TD
    A["Layer 0 vision injection
(tone anchor: East Asian rural village)"] --> B["Layer 1 rulebook
(generates objective/reward slot skeleton)"] B --> C["L2 AI assist
(fills NPC motive, one-line background)"] C --> D{"Layer 4 review gate
(tone/consistency lint)"} D -->|pass| E["Layer 3 data finalized
(rulebook computes reward values)"] D -->|reject| B E --> F["Merged into build"] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; class B,D code; class C ai; class E data; class F pass;

Let me walk this flow once as a worked transcript. Say we are making one odd-job quest for a rural village.

Step 1 — the Layer 1 rulebook draws the skeleton. The rulebook is deterministic. It comes out with nothing but the slots filled.

[Quest skeleton — rulebook output]
Region: 향촌_3
Requester role: herbalist
Objective type: collect (medicinal herbs x6)
Obstacle: wolf spawn (night)
Reward slot: { gold: ?, exp: ?, item: gathering sickle }

Step 2 — ask the AI for the details. Here is the full prompt. (The Layer 0 vision is fed in first, as context.)

You are a quest writer for an MMORPG set in an East Asian fantasy rural village. Tone anchor: "humble countryside, with superstition and fear underneath." Take the skeleton below and write a one-line motive for the requester, an NPC name, and the two-sentence request shown to the player. Never make up reward values (the system decides them). Output JSON. Skeleton: { region: 향촌_3, requester: herbalist, objective: collect 6 medicinal herbs, obstacle: wolves at night, reward item: gathering sickle }

Step 3 — Claude's raw output. (Reproduced as-is, not summarized.)

{
  "npc_name": "삼월 어멈",
  "motive": "With wolves howling every night her herb patch is being ruined, and she cannot bring herself to go dig up the last of the herbs.",
  "quest_text": "When night falls, the wolves tear up my herb patch. I am too frightened to set foot out there, so please dig up just six roots of medicinal herb at the foot of the mountain in my stead.",
  "reward": { "gold": 120, "exp": 300 }
}

Stop here for a moment. The AI added a reward field without being asked to. This is a live demonstration of why the first and second axes have to stay separate. Reward values (Layer 3) are the L1 rulebook's seat, not the AI's (L2). Hand them to the AI and the numbers wobble on every call, and the reward curve collapses.

Step 4 — human verification and veto. The reviewer does two things. (1) Deletes the reward field — that is a cell for the rulebook to fill. (2) Checks the tone. The NPC name "삼월 어멈" (Granny Samwol), the one-line motive — wolves wreck her herb patch every night, so she cannot bring herself to go dig the last herbs — and the two-sentence player-facing request all fit the rural-village tone. Pass. If the AI had inserted an out-of-setting phrase like "a commission from the mages' guild," it would be rejected here and sent back to the skeleton stage.

Step 5 — Layer 3 data finalized. The rulebook refills the deleted reward slot. It is a deterministic formula tied to region level and objective difficulty. gold: 85, exp: 240. Not the 120 and 300 the AI emitted arbitrarily — values that sit on the curve.

This one lap is the standard cycle of the center-of-gravity cell. Rulebook for the skeleton, AI for the flesh, a human for the gate, and the rulebook again for the numbers. All 1,000 pieces of content run this same cycle. Because the cell is fixed, nobody re-litigates "who makes this" every time.


6.1.6 Five Questions for Choosing a Cell

Five questions help decide which cell of the grid a new piece of content belongs in. Write them down and answer them together every time a mass-production item comes up in a meeting, and cell placement becomes consistent within a quarter.

One: how heavy is the production load? How many do you need by launch? If N is over 100, the L0 row is close to impossible.

Two: how strong is the consistency requirement? If cross-content consistency is the core of the experience, the review gate (Layer 4) has to be strong; if variety is the core, there is room to climb to a higher automation row.

Three: can you tolerate non-determinism? Is this a domain where a slightly different result every time creates richness, or one where the same result is the core of trust?

Four: what does review cost? Whether it is 5 minutes or 30 minutes per piece sets the length of the operating cycle.

Five: what does an incident cost? Can you freely scrap and rewrite, or does one slip ship straight into a player-facing incident?

Throw these five at side quests and the answers converge in one direction: more than 1,000 needed (L0 impossible), consistency requirement lower than main quests, non-determinism acceptable, review 5–10 minutes, incident cost low (each one can be scrapped individually). With all five pointing the same way, the L2 row × Layer 2 cell is the natural fit. Throw the same five at main quests and they converge the opposite way: 50 quests, the highest demands on consistency and narrative depth, non-determinism unacceptable, review cost high, incident cost very high — that is the L0 cell.


6.1.7 Four Common Pitfalls

Even with the grid drawn, teams fall into the same traps. Four of them repeat.

First, starting from the L3 row. Set out with the expectation that "AI will just handle all 100" and review explodes. Settle the L1 cell first, climb to L2, and apply L3 carefully, to a small subset only. In the mini pipeline above, the single motion of a human deleting the reward field is a small demonstration of why the L3 row is dangerous.

Second, delegating to AI wholesale, with no rulebook. "Make me 100 side quests" summons the generic-RPG average. Only when the Layer 1 rulebook sets the skeleton first and the AI fills in the flesh on top of it do you get your game's content. Writing a rulebook is the most laborious and least fun job in PCG, but skip it and everything mass-produced above it sinks down to the average.

Third, an empty review gate (Layer 4). When AI output flows into the build automatically, consistency incidents follow directly. Whatever the cell, a human gate is mandatory.

Fourth, choosing tools on cost alone. LLM API prices drop every quarter; the cost of a consistency incident does not. A tool decision should weigh API cost plus the combined cost of consistency and review time.


6.1.8 Measurement — Six Months After Moving to the Center-of-Gravity Cell

On Project A, we moved side quests from the L0 cell to the L2 cell and measured the six months that followed. In the numbers below, the absolute values are the author's estimate (unverified); what was observed in actual measurement is the direction and ratio of the change.

Item L0 period After the L2 move
One quest per writer \~4 hours \~50 minutes (30 min metadata + 5 min AI + 8 min review)
Weekly output 5 30–40
Scrap rate nearly 0% \~20%
Consistency incidents (per quarter) 3–5 5–8 (normal after reinforcement)
Writer satisfaction (out of 10) 8 6 → 7 (after the policy fix)

The scrap rate rose to 20%, but with production speed at 6–8x, net throughput grew 4–5x. Consistency incidents rose slightly, to 5–8 per quarter, but with the review gate and the rulebook reinforced, they returned to the normal range within the quarter.

The biggest change was not in the numbers but in the people. At first the writers said they felt like "reviewers on an assembly line," and satisfaction dropped from 8 to 6. To recover it, we added a policy that explicitly guarantees writer time on main quests and on signature side quests (1–2 per city). We pinned it down so that the production line would not be a thing that sucks up writer time, but a tool that sends that time back to the main content. Six months later, satisfaction was back at 7.

One thing to take from this measurement: a decision to move cells has to carry throughput, writer-time allocation, and satisfaction together. Watch throughput alone and the production numbers succeed — and the people leave.


6.1.9 Layer Decomposition First, PCG on Top

The general thesis — that Layer decomposition is the precondition for procedural generation — is in §2.3.6. Here I only look at how it shows up on the PCG grid. On a team where the horizontal axis (Layer 0–4) is blurry, no cell runs stably. If nobody knows where the Layer 0 vision lives, every generator runs with an empty tone anchor and produces the generic-RPG average. If the Layer 1 rulebook and the Layer 2 body are mixed in one file, fixing one line of a rule means touching dozens of places in the body at the same time. If Layer 3 data is typed into the body, one reward-curve adjustment costs a writer a week — this is why the mini pipeline above keeps rewards in a separate slot.

So what to check before adopting PCG is not tool choice but whether the horizontal axis is decomposed. On a team with the five layers in place, attaching an L1 generator costs one writer one quarter. On a team where the five layers are tangled, the same adoption gets scrapped within two quarters on consistency incidents.

The five layers do not need to be perfect from day one. Separate gradually, keep the interfaces narrow. In the first quarter, splitting off just one Layer 0 tone line and one Layer 1 rulebook opens a seat for a generator to come in. That is not license to postpone forever, though. If the Layer 2 body and the Layer 3 data stay one lump to the end, even the concrete tool of the next chapter will have nowhere to sit.


6.1.10 Next Chapter Preview

The next chapter dissects one concrete tool that occupies the grid's center-of-gravity cell: proj_city_hunting_generator, which mass-produces per-city hunting grounds. We will see how input metadata, the rulebook skeleton, the AI body, and the verification gate are tied into a single cycle — and how this chapter's mini pipeline scales up to the size of a real tool.


Key Takeaways


Try It Yourself — Putting One Piece of Content on One Cell

setup. Pick one candidate for mass production (for example, side quests). Split the Layer 0 tone line and the Layer 1 rulebook skeleton (slot definitions) into separate files. Leave the reward-value slot empty on the rulebook side.

prompt. Feed the vision in as context, then give the skeleton, and state explicitly: "do not make up reward values; output JSON." You can adapt the Step 2 prompt above as-is.

verify. Check three things. (1) If the AI inserted a reward field on its own, delete it (L3 is the rulebook's seat). (2) If there are out-of-setting words, reject back to the skeleton stage. (3) Only for what passes, let the rulebook fill in the reward values and put it in the build.

Solo Scale-Down. You do not need a team. Write one rulebook (five slots) and one tone line yourself, as plain text files. Run ten quests through the cycle above and count how many you reject in review. If the rejection rate goes past 30%, the cell is wrong — tighten the rulebook skeleton, or drop one row down (L1) and look again. When the rejection rate stabilizes, that is the signal that the cell works at your scale.

6.2 city_hunting_generator — 30 Cities in 4 Weeks

Primary readers: MMORPG designers responsible for content production at scale (mid-size team, 10–50 people) Scaled-down version for solo/hobbyist readers: §6.2.10, "If You're Solo, Just This Much"

I still remember the math I did the day the schedule first landed on my desk: 30 cities needed by launch. One city consists of a 5–10 line introduction, 3–5 hunting grounds, 5–10 NPCs and 2–3 side quests per hunting ground, 1–3 specialty items, and 1 city boss. Crafting one city by hand takes 1–2 weeks. At 30 cities, that is a single writer spending 6 solid months on nothing but cities.

We did not have those 6 months. The writers' time was tied up in the main quest and the signature characters, and the 30 cities had to be built in parallel with that work. My first impulse — "why not just ask the AI to make 30 cities?" — collapsed quickly. Hand over the whole job and you get 30 fantasy towns that all look alike. Instead of that impulse, I built a tool, city_hunting_generator. This chapter looks at how it ties input, rulebook, AI, and verification into a single cycle — and what actually comes out, and what gets discarded, when you run that cycle all the way through once.

Author's operational note The city_hunting_generator in this chapter is an anonymized version of a real tool I run in the company R&D folder. The file names, code structure, and verification items faithfully follow the real tool; city names (silvermark and so on) and company-specific names were replaced for the book. The output text is a reconstruction of actual sessions.


6.2.1 Humans Handle Only the Metadata and the Final Review

The tool runs in four stages. The key point: stages 1 and 3 are deterministic (the rulebook), and only stage 2 is AI. With the rulebook holding both the skeleton and the verification in place, the AI sandwiched in the middle can give slightly different answers every time without shaking cross-city consistency. Humans enter only at the first input (metadata) and the final gate (review).

flowchart TB
    A["Input: city metadata yaml
(human: 15~20 min/city)
3 lore_seeds · forbidden_names auto-attached"] A --> B["Stage 1 deterministic: rules.py
generate_skeleton()
hunting ground count · enemy distribution · reward curve · boss"] B --> C["Stage 2 AI: 3 prompt types
L0 vision cached + L1 rules injected
+ L2 neighboring city text
→ intro · NPCs · side quests"] C --> D{"Stage 3 deterministic: lint
name dupes · forbidden words · voice
· length · reward range"} D -->|violation alert| E["Stage 4 writer review gate
(5~10 min/city)
discard/regenerate decisions"] E -->|re-request| C E -->|pass| F["Merged into build"] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; class B,D code; class C ai; class A,E human; class F pass;

In this diagram, human hands touch only two places: at the top, where one clean page of metadata goes in, and at the bottom, where the tone and narrative judgments lint cannot catch are made. The tedious skeleton generation and bulk text production in between are run by the rulebook and the AI. The decisive design choice is that even when lint (stage 3) finds a violation, it does not auto-discard — it only raises an alert to the writer gate (stage 4). The reason is in §6.2.5.


6.2.2 Input — One Page of City Metadata

The writer writes one page of metadata per city. It takes 15–20 minutes. Short, but this one page is the entire input for the next three stages.

# city_021_silvermark.meta.yaml
city_id: city_021_silvermark
region: west
climate: cold_arid
dominant_faction: scholar_guild
cultural_tone: scholarly_strict
level_range: [25, 30]
lore_seeds:
  - Was the center of a magic seal 100 years ago
  - The first sign of the seal weakening was observed in this city
  - Site of the scholar guild headquarters
neighbors: [city_018, city_023]
# forbidden_names: (auto-attached by script — no writer input needed)

(The three lore_seeds read: "was the center of a magic seal 100 years ago," "the first sign of the seal weakening was observed in this city," "site of the scholar guild headquarters." The final comment notes that forbidden_names is attached automatically by a script — no writer input needed.)

The most important slot is lore_seeds. 3–5 key events anchor the city's identity. Too few and the AI spits out a generic fantasy city; too many and the events contradict each other. In my experience, 3 was the most stable.

forbidden_names is not filled in by the writer. A script reads the list of existing city and character names and attaches it to the metadata automatically. Once 30 cities × roughly 50 NPCs each pile up, checking 1,500 names for duplicates in your head is impossible. There is no need to hand-write "make sure it doesn't overlap with other cities' NPCs" every single time.


6.2.3 Stage 1, the Rulebook — Building the Skeleton Deterministically

The rulebook takes the metadata and builds the city's structural skeleton. The code is simple.

# city_hunting_generator/rules.py (skeleton)
def generate_skeleton(meta):
    region_rules = REGION_RULES[meta.region]
    hg_count = region_rules.hunting_grounds_range.sample()
    enemy_dist = ENEMY_RULES[meta.climate][meta.dominant_faction]

    skeleton = {
        "hunting_grounds": [
            {
                "id": f"{meta.city_id}_hg_{i}",
                "level": meta.level_range[0] + i,
                "enemy_types": enemy_dist.sample(k=3),
                "reward_curve": calc_reward(meta.level_range[0] + i),
                "npc_count": region_rules.npc_per_hg,
                "sidequest_count": region_rules.sidequest_per_hg,
            }
            for i in range(hg_count)
        ],
        "boss": {
            "id": f"{meta.city_id}_boss",
            "level": meta.level_range[1] + 2,
            "pattern": BOSS_PATTERNS[meta.region],
        },
    }
    return skeleton

The result is deterministic. Same metadata in, same skeleton out. Whether the reward curve sits within the standard range for each region and level, and whether the enemy distribution matches the climate and faction rules, is guaranteed in code and covered by regression tests. This stage is never handed to the AI. If the AI pulled a different reward number on every call, cross-city balance would wobble on the spot.

Feed in silvermark's metadata and rules.py returns an empty skeleton: 4 hunting grounds (city_021_silvermark_hg_0 through hg_3), 6 NPC slots and 3 side quest slots per hunting ground, and 1 boss at level 32. It is a table of cells to fill — no names, no text yet. Filling those cells is the job of stage 2, the AI.


6.2.4 Stage 2, the AI — Generating Natural-Language Text

After the rulebook builds the skeleton, the AI fills in natural-language text on top of it. This is where the city introduction, the NPC names, appearances, and short backstories, the side quest synopses, and the specialty item flavor text come from.

The call pattern follows the four-layer context injection structure exactly. Cache the L0 vision (world_premise + tone_manifesto), selectively inject the L1 rules (city_naming_rule + region_west_lore), add the L2 adjacent text (NPC name lists from other cities), and attach the task instruction at the end. The city introduction prompt is in a form you can copy and use as is.

[L0 context] world_premise + narrative_pillar + tone_manifesto  (cached)
[L1 context] city_naming_rule, region_west_lore
[Input] city_021_silvermark.meta.yaml + 3 lore_seeds

Write a 6~8 line introduction for this city. Weave all three lore_seeds in naturally,
and cut RPG stock phrases like "a peaceful village." Keep the tone scholarly and strict, sentiment restrained.
Output the text only, no preamble or commentary.

(The instruction, in brief: write a 6–8 line introduction for this city; weave all three lore_seeds in naturally; cut RPG stock phrases like "a peaceful village"; keep the tone scholarly and strict, sentiment restrained; output the text only, no preamble or commentary.)

The same pattern repeats as is for the NPC batch and the side quest calls. Only the context and the output format differ. The NPC call carries the forbidden_names slot (no duplicating NPC names from other cities); the side quest call carries the constraint "pick the quest-giving NPC from the NPCs generated above."

So what actually comes out? Let's run it through once, end to end.


6.2.5 One Cycle, All the Way Through — Output, Verification, Discard

Writing only the abstraction "the AI fills in the text" tells you nothing about what this tool actually spits out. Let's follow one full cycle for the city of silvermark, from input to discard.

[Stage 2 AI output — city introduction] (reconstructed from an actual session)

Silvermark is a city of the seal. 100 years ago, seven scholars raised a sealing circle beneath this ground to hold back the flood of mana from the west, and the city grew over that circle, one stone at a time. The bell tower of the scholar guild headquarters is an observatory that reads the seal's vein, not a spire raised toward any god. For the past three years, one line has repeated in the observation logs: "The vein is cooling." The first sign of the seal's weakening trembled first on this city's gauges, nowhere else. The scholars in the streets do not say it aloud. Only the nights when the bell tower's fire never goes out have grown more frequent.

All 3 lore_seeds made it into the text (seal epicenter, first sign of weakening, scholar guild headquarters); there is no "peaceful village that welcomes adventurers" stock phrase; and the scholarly_strict tone holds (sentiment restrained, a vocabulary of observation and record). This is what a passable output looks like.

Next, the same cycle fills 3 of the 6 NPC slots in hunting ground hg_0.

[Stage 2 AI output — 3 NPCs] (reconstructed from an actual session)

1) Doren Vale — Observation assistant at the bell tower. In his 50s, a brass loupe permanently fixed over one eye. Has recorded the fine tremors of the seal's vein for 30 years. A man of few words; converses only in numbers. "Today, 12.4. Yesterday, 12.1. It's climbing. That is not a good thing."

2) Mira Kost — Archivist of the guild's document vault. In her 30s, ink stains on her fingers that never wash off. She guards the original blueprints of the sealing circle, yet believes every scholar who could actually read them is dead. Deeply wary of outsiders.

3) Grem — Furnace keeper beneath the bell tower. Identity unknown, age unknown. His only job is keeping the bell tower's fire from going out, and to anyone who asks about the seal he answers only, "Just watch the fire." (flagged as uncertain — self-reported by the AI)

Note that the AI attached an uncertainty flag to the third NPC, Grem, on its own. A good prompt makes it possible for the AI to say, "I'm not confident about this one." Now stage 3 lint runs against this output bundle.

[Stage 3 lint output] (actual format)

[PASS] Length check: introduction 7 lines (standard 6~8)
[PASS] Reward range: hg_0~hg_3 reward_curve within standard range
[WARN] Name duplication: "Mira Kost" — vs. "Mira Veldt" of city_014_riverhold,
       surname (Kost/Veldt) differs but given name (Mira) identical. Near collision in forbidden_names.
[PASS] Forbidden vocabulary: tone_manifesto violations 0
[WARN] Voice consistency: "Grem" dialogue voice_lint confidence 0.62 (below 0.70 threshold)

(The five lines read: PASS on length — introduction is 7 lines against a 6\~8 standard; PASS on reward range — hg_0 through hg_3 reward_curve within standard range; WARN on name duplication — "Mira Kost" vs. city_014_riverhold's "Mira Veldt," different surnames but the same given name Mira, a near collision in forbidden_names; PASS on forbidden vocabulary — 0 tone_manifesto violations; WARN on voice consistency — Grem's dialogue at voice_lint confidence 0.62, below the 0.70 threshold.)

Lint caught 2 violations but auto-discarded neither. It only raised them as WARN to the writer gate. This is the design core previewed in §6.2.1. Give the verifier the power of automatic rejection, and within a month or two the writers will flip that switch off. The machine kills intended variation along with the real violations, and it robs the writers of the chance to gauge that boundary for themselves. So the machine is in charge of picking out suspicious candidates, but the final call on whether a candidate lives or dies stays in human hands.

[Stage 4 writer review — verdicts and discard]

The writer handled the 2 alerts like this.

Once the writer decides to discard, one re-request goes around: "Discard the Grem slot. Regenerate a furnace keeper NPC that fits the same hunting ground's scholar guild tone (observation, records, rigor). No mysticism vocabulary." The AI answered with an old man who logs the temperature of the bell tower furnace — a man who sees even fire as data — and that output passed with voice_lint 0.81. The cycle of input → skeleton → text → verification → discard → regeneration closes here.

This one full lap is the Show standard for this entire book. Unless you watch, at least once and all the way through, what the tool spits out, what gets caught, and what a human kills, the sentence "we mass-produced it with AI" is hollow.


6.2.6 The Discard Rate Is Not Tool Failure but a Signal from the Gate

In the cycle above, 1 NPC was discarded. Across a whole city, the discards pile up further. Review time averages 5–10 minutes per city; the discard rate is about 20% for NPCs and about 33% for side quests.

Let me be honest about where these rates come from. The discard rates were counted by hand while personally reviewing 5 cities, silvermark included, in the early adoption period. For NPCs, 6 of 30 reviewed were discarded (20%); for side quests, 5 of 15 (33%). With a sample of only 5 cities, the right way to read these is not as precise population rates but as directional values — "one in five, one in three." The cumulative rate after all 30 cities are reviewed could come in lower, or higher depending on the character of the hunting grounds.

What matters is that a 0% discard rate is not the goal. Zero discards is closer to a signal that the review has become a formality. When one in five NPCs gets discarded for the wrong tone, and one in three side quests gets regenerated for failing to connect to the lore_seeds, the review gate is actually working.


6.2.7 Measurement — 30 Cities in 4–5 Weeks

Before and after the tool. The time figures below are measured averages from the early cities, silvermark included; the "before" column is the writers' estimate from the hand-crafting era before the tool. None of the numbers have been doctored.

Item Before (by hand) After (measured)
Time to write 1 city 1–2 weeks About 30 min (metadata 15 + AI 5 + review 8)
Total span for 30 cities 6 months of one writer 4–5 weeks
Discard rate (NPCs) — (all written directly) About 20% (6 of 30)
Discard rate (side quests) About 33% (5 of 15)
Consistency incidents (per city) Nearly none 0–1

The table makes it look like the numbers are the whole story, but the real effect came from a different cell. With writer time freed from city production, one writer could substantially increase main quest output per quarter (the exact multiple varies by quarter, so I won't pin it down — the direction is "main output clearly went up"). The production tool worked as a tool that releases writer time rather than absorbing it (the warning from §6.1.8 applies as is: if the writers feel they have been turned into "review machines," the tool gets rejected).


6.2.8 Content That Doesn't Go into the Generator

Even with automation covering more ground, the following stays outside the tool.

Content Why it stays outside the tool
Main quest text Consistency and narrative depth tie directly into the game's identity
Boss patterns and staging Heavy on visual and interaction detail — a designer's hands are faster
Signature main characters Requires a full voice_profile; cannot be mass-produced
Branch endings The writer's direct decision territory
1–2 signature side quests per city The writer picks them and builds them personally

The fact that something can be mass-produced must not auto-convert into the decision that it should be. As the silvermark cycle showed, the tool produces 5 out of 6 NPCs well. But the one signature NPC who carries the city's core tension — "the seal is cooling" — the writer shapes by hand. When the boundary of automation is clear, the production tool becomes the very thing that protects that core territory.


6.2.9 Five Common Failures

Failure pattern Why it fails Fix
Writing only 1–2 lore_seeds AI output flattens to the generic RPG average Require 3 or more (§6.2.2)
Asking the AI to mass-produce wholesale, no rulebook "Make me 30 cities" → 30 near-identical towns Stage 1 rulebook cannot be skipped (§6.2.3)
Relying on writer review alone, no lint Reviewers burn their time on trivial rule violations Run automated first-pass verification first (§6.2.5)
Skipping the name duplication check 1,500 names cannot be deduplicated in anyone's head Auto-attach forbidden_names (§6.2.2)
Not measuring writer satisfaction Throughput rises, but steal the writers' time and the tool gets rejected Explicitly guarantee main content time (§6.2.7)

The fifth is the one most often missed. For a writer to enjoy making calls like discarding silvermark's Grem, they need time left over to craft things themselves — not just review the mass-produced output. Measure throughput only and skip measuring writer time, and the tool succeeds on the KPIs while the people leave.


6.2.10 Try It Yourself — One Step You Can Take Today

If you're solo, just this much: You don't need rulebook code. Pick one city or region from your own game (or a game you love), hand-write the metadata in the §6.2.2 format (the 3 lore_seeds are the heart of it), and run the introduction prompt from §6.2.4 once, pasted as is. Then pick the one NPC whose tone feels off and push back: "This NPC clashes with the city's tone — discard and regenerate." That is when the review gate stops being a concept and becomes, in your own hands, the bundle of judgments it really is.

If you're on a team, start with this one step. Build the one-page metadata yaml template and the forbidden_names auto-attach script first. The rulebook skeleton (generate_skeleton) and lint come after. With just the input template and the name duplication check, you can already head off the two most common failures that collapse AI text production into "30 near-identical towns."


6.2.11 A Preview of the Next Chapter

6.3 covers the NPC Persona/Squad pipeline. Where 6.2's generator mass-produces NPCs like Doren and Mira individually, Persona/Squad binds those NPCs into groups. It is how you make the five NPCs of one hunting ground work as a small society rather than a collection of unrelated dolls.


Key Takeaways

Next Chapter Preview

6.3 NPC Persona and Squad — From a Mannequin Museum to a Small Society

Primary readers: MMORPG designers who own NPC and hunting-ground content (mid-size teams of 10–50) Scaled-down version for solo/hobbyist readers: §6.3.10 "If You're Solo, Just This Much"

I remember the day I used the 6.2 generator to mass-produce five NPCs for one hunting ground and put them in the game. Names, appearances, short backstories — all filled in. I placed them by dropping coordinates. But when I actually walked through that hunting ground, it felt strangely dead. Five people shared the same space and never once mentioned each other. Two of them stood overlapping on the same rock. Someone needed to play the merchant, but all five were scholars. Doren and Mira were each perfectly fine NPCs on their own; bundled together, they were a collection of mannequins.

This is the mannequin-museum state. Every individual NPC is built, but the group is not alive. This chapter covers the pipeline that binds those five into a small society. The core decomposition is Persona and Squad. To use an office analogy: a Persona is an employee's business card, and a Squad is one team's org chart. Stack fifty business cards with no org chart and the company does not run. And the spine of this chapter is the final stage — running one full cycle with the AI to verify that the bundled group talks and moves like people who actually know each other.

Author's Operating Note The Squad pipeline in this chapter is an anonymized version of the NPC Persona/Squad tool I run in my company's R&D folder. The yaml structure, the checks, and the voice_lint thresholds faithfully mirror the real tool; the city and NPC names are swapped for the book, the same as in 6.2. The output text is reconstructed from real sessions.


6.3.1 Persona Is the Business Card, Squad Is the Org Chart

A Persona is an individual NPC's identity. It holds the name, appearance, voice_profile, and role. What the 6.2 generator produces is Personas. Doren Vale and Mira Kost are each one Persona.

A Squad is the unit that binds those Personas into a group. It defines how five people are distributed across roles in one hunting ground, how they relate to each other, and how they move.

Unit What it holds Who makes it
Persona name, appearance, voice_profile, role generator (6.2)
Squad role distribution, relationships, movement Squad pipeline (this chapter)

If you do not separate the two, you get blocked in two ways at once. Mass-produce only Personas and you get a mannequin museum; try to build Squads first and you have no Personas to fill them with. Separating them keeps each unit simple to operate. Separation is not severance, though. The point is to lay reuse and verification paths between the two units, and that is the main business of this chapter.

This Persona→Squad decomposition is not mere housekeeping; it opens a longer road. Only when NPC groups are formalized into roles, relationships, and numbers can you later reach dynamic reactivity — where world state (accumulated player behavior) perturbs NPC values and those values become quest trigger conditions. This chapter only points at the entrance to that progressive application; what it covers head-on stops at conservative mass production gated by human review.


6.3.2 Input — One Page of Squad Metadata

The Squad skeleton starts from one page of metadata per hunting ground. Same philosophy as the city metadata in 6.2: a human pins down only the role distribution and the intended relationships, and the rulebook and the AI do the filling.

# city_021_hg_3.squad.yaml
squad_id: city_021_hg_3_squad
hunting_ground: city_021_silvermark_hg_3
type: hunting_ground_residents
size: 5
roles:
  - role: quest_giver
    count: 1
    voice_traits: [authoritative, scholarly]
  - role: lore_keeper
    count: 1
    voice_traits: [scholarly, withdrawn]
  - role: merchant
    count: 1
    voice_traits: [practical, dry]
  - role: bystander
    count: 2
    voice_traits: [varied]
relationships:
  - between: [quest_giver, lore_keeper]
    type: mentor_and_former_student
  - between: [merchant, bystander_1]
    type: regular_customer
movement_pattern: stationary_with_shifts

The most important slot is relationships. With zero relationships, the five stay strangers forever. With too many (five or more for a five-person squad), players have too much to keep track of and the relationships drown each other out. In my experience, 2–3 core relationships per five-person Squad is the most stable. voice_traits is the device that gives the five distinct voices. Fill all five with scholarly and the verification stage flags it as voice homogenization.


6.3.3 Stages 1 and 2 — Rulebook Skeleton and Persona Fill

The rulebook sets the standard for the Squad skeleton first. Default values for size, role distribution, relationship density, and movement pattern are coded in per hunting-ground region and type.

# npc_squad/templates.py (excerpt)
SQUAD_TEMPLATES = {
    ("west", "hunting_ground_residents"): {
        "size_range": (4, 6),
        "role_distribution": {
            "quest_giver": 1,
            "merchant": 1,
            "lore_keeper": (0, 1),
            "bystander": (1, 3),
        },
        "relationship_density": 2,        # recommended relationship count
        "movement_pattern": "stationary_with_shifts",
    },
    ("east", "outpost_squad"): {
        "size_range": (3, 4),
        "role_distribution": {
            "commander": 1,
            "scout": 1,
            "support": (1, 2),
        },
        "relationship_density": 1,
        "movement_pattern": "patrol_loop",
    },
}

This stage is deterministic. For a western-residents Squad, the accident of five quest_givers and nothing else is impossible at the code level. If a role distribution steps outside the rules, it gets blocked on the spot.

Next, each slot is filled with a Persona. There are three paths. If the pool has a Persona that fits, reuse it (appearance weight +1); if not, generate a new one with the 6.2 generator; and if it is a key figure in the main quest, a writer writes it by hand. For silvermark's hg_3 Squad, the quest_giver and lore_keeper slots were filled with Mira and Doren, already mass-produced in 6.2, and the merchant and two bystanders were newly generated. Up to this point it is the same cycle as the 6.2 generator. The real work of this chapter comes next: verifying that the bundled group actually behaves like a group.


6.3.4 One Cycle, End to End — Relationship Fill-In, Movement, and Consistency Checks

If I only wrote, abstractly, that "the AI fills in the relationships," you would have no idea what this pipeline spits out. So let us follow the back half of one cycle for the silvermark hg_3 Squad, end to end — from generating the relationship text to discarding and re-requesting.

Stage 3 — AI Relationship Fill-In

The relationship tag entered in the Squad skeleton (mentor_and_former_student) is an abstraction; it is invisible in the game. Stage 3 turns it into a one-line description to be planted in NPC dialogue and events. The prompt is in copy-paste-ready form: it hands the AI the L0/L1 context, the two Personas, and the relationship tag, then asks for a one-to-two-line background usable in game dialogue — the two speech styles must not clash, scholarly-strict tone, no mysticism, no "old friends" clichés, body text only.

[L0 context] world_premise + tone_manifesto  (cached)
[L1 context] city_021_silvermark.lore (ruled by the scholars' guild, scholarly_strict)
[Persona 1] quest_giver — Mira Kost, guild archive librarian, 30s, ink stains
[Persona 2] lore_keeper — Doren Vale, bell-tower observation assistant, 50s, speaks only in numbers
[Relationship tag] mentor_and_former_student

Describe the relationship between these two (mentor and former student) in just 1~2 lines,
as background usable in game dialogue. Doren is numbers, Mira is documents — keep the two
speech styles from clashing. Strict scholarly tone, no mysticism, no clichés like "old friends". Body text only.

[Stage 3 AI Output — Relationship One-Liner] (reconstructed from a real session)

Twenty years ago, Doren taught Mira the notation used in the seal-array observation records. Now their positions are reversed: Mira copies the readings Doren takes into the archive ledgers. Every Tuesday the two of them argue, briefly, over the one cell where the readings and the ledger disagree.

This output is good. mentor_and_former_student became concrete; Doren's "numbers" and Mira's "documents" were tied into one scene (copying readings into a ledger) without clashing; and the scholarly_strict tone held. The same prompt is repeated for the merchant–bystander_1 regular_customer relationship.

Stage 4 — Movement Synthesis

An NPC standing in one spot all day becomes a mannequin again. The rulebook fills in the movement pattern. stationary is a fixed position (guards, bosses); stationary_with_shifts nudges the position every 8 hours (the common case); routine_loop runs on a timetable (residents); event_driven moves only on triggers (quest NPCs). This is deterministic, so the AI is not called.

Stage 5 — Squad Consistency Lint (This Pipeline's Gate)

Now we test whether the five, bundled, actually behave like a group. Where the 6.2 lint looked at individual NPCs, this lint looks at group consistency.

[Stage 5 Squad Lint Output] (actual format)

[PASS] Role distribution: quest_giver 1 · lore_keeper 1 · merchant 1 · bystander 2 (rule satisfied)
[PASS] Relationship density: 2 (recommended 2, satisfied)
[WARN] Voice diversity: scholarly family 3/5 — quest_giver·lore_keeper·bystander_2
       have voice_profile cosine similarity 0.83 (exceeds 0.80 threshold). Homogenization risk.
[WARN] Movement collision: merchant·bystander_1 coordinates overlap within 1.5m radius, 14:00~16:00 window
[FAIL] Relationship visibility: 2 relationships defined, but 0 mentions of any other member in the 5 NPCs' dialogue.
       Relationships exist only in data — in-game visibility 0.

The lint caught three items: a voice-diversity WARN (3 of 5 voices in the scholarly family, cosine similarity 0.83 against a 0.80 threshold), a movement-collision WARN (merchant and bystander_1 overlapping within a 1.5m radius in the 14:00–16:00 window), and a relationship-visibility FAIL (2 relationships defined, 0 mentions of any other member in the five NPCs' dialogue — the relationships exist only in data). None of the three is discarded automatically; all three go up to the gate — the machine surfaces the suspects, but a human decides what lives and what dies. Same design as §6.2.5.

[Stage 6 Writer Review — Verdicts and Discards]

The writer handled the three alerts like this.

Two of the three closed via a rule fix or regeneration; the last FAIL is the heart of this pipeline. The writer requested one additional dialogue-branch line for Doren — a single line in which the relationship with Mira slips out in passing, an aside rather than exposition, in a scholar's tone.

Insert just one line into Doren's dialogue where the relationship with Mira slips out in passing.
Not expository — as an aside. Scholarly tone, one line of dialogue only.

[Re-request Output]

"That diagram is in the archive. Ask Mira. ...Twenty years ago I was the one teaching her how to read them. These days it's the other way around."

The moment this one line goes in, the relationship between the two NPCs moves from the data sheet onto the game screen. Input (Squad metadata) → skeleton → Persona fill → relationship fill-in → movement → consistency check → discard and visibility decisions: one full cycle closes here.

This one loop is this chapter's standard of Show. The sentence "we bound the NPCs into a society with Squads" is hollow unless you have watched, at least once, a human close a relationship-visibility-0 FAIL with a single line of dialogue.


6.3.5 The Full Persona→Squad Flow

Here is the cycle above in a single diagram. The key points: stages 1, 2, 4, and 5 are deterministic (rulebook and lint), only stage 3 is AI, and human hands touch only the input at the top and the gate at the bottom.

flowchart TB
    P["Persona pool
(6.2 generator output)
Doren, Mira, ..."] --> FILL A["Input: Squad metadata
(human, 10–15 min per hunting ground)
role distribution + intent for 2–3 relationships"] A --> B["Stage 1, deterministic: templates.py
standards for role distribution, relationship density, movement"] B --> FILL["Stage 2: Persona fill
reuse / generator / writer-authored"] FILL --> C["Stage 3, AI: one-line relationship fill-in
L0 caching + voice_traits conflict check"] C --> D["Stage 4, deterministic: movement synthesis
shift, routine, event_driven"] D --> E{"Stage 5, deterministic: Squad lint
role distribution, voice diversity,
movement collisions, relationship visibility"} E -->|"alert"| F["Stage 6: writer review gate
(5–10 min per hunting ground)
decide: discard, rule fix, relationship exposure"] F -->|"re-request"| C F -->|"pass"| G["Into the build"] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; class B,D,E code; class C ai; class A,F human; class P data; class G pass;

Human hands touch only two places: the spot at the top where role and relationship intent is set, and the spot at the bottom where someone judges the tone and narrative that lint cannot catch. In between, the rulebook runs the skeleton, the movement, and the checks, and the AI writes the relationship text.


6.3.6 Three Devices That Make Relationships Visible in the Game

The 관계 노출 (relationship visibility) item is the most frequent FAIL in the stage 5 lint, because relationships tend to live only in the data. There are three devices for pulling a relationship into the game.

First, dialogue mentions. As with Doren's line in §6.3.4, an NPC mentions another member in one line. Cheapest, and the most effective.

Second, route crossings. Every Tuesday, Doren and Mira can be observed together in the archive, inside the game. A player who happens to see it thinks, "are those two connected?" If the stage 4 movement agrees with the relationship, this comes out naturally.

Third, branch conditions. Refuse the quest_giver's request and the lore_keeper's affinity drops with it. This third device is the entrance to the progressive application mentioned in §6.3.1 — the point where a relationship goes beyond mere description and starts affecting game state.

You do not need all three. In a five-person Squad, exposing just the 2–3 core relationships through the first and second devices already changes how the hunting ground feels. Overdo it and players have too much to memorize. At review, the writer picks which relationships to expose and leaves the rest as data.


6.3.7 Measurement — Honestly

I compare before and after adopting the tool. No fabricated figures. The times and rates are values I counted while personally reviewing the first several hunting grounds, including silvermark; the "before" column is a writer's estimate from the handcraft era.

Item Before (manual, estimated) After (measured)
Bundling one hunting ground's Squad about 3–4 hours about 25 minutes (12 min metadata + 5 min AI + 8 min review)
Relationships exposed (dialogue/routes) 0–1 per hunting ground 1–2 exposed out of 2–3 core
Movement collisions (2+ NPCs at same coordinates) 2–3 per hunting ground blocked up front by lint, 0–1
Voice-homogenization discards — (no check) 0–1 of 5 regenerated

The sample is small — a handful of hunting grounds — so read these as directional values, not precise population rates. The biggest change does not fit in the table. Because the lint's 관계 노출 FAIL forcibly shoves "this hunting ground shows zero relationships" in the writer's face, mass-produced content shipping as a mannequin museum became structurally rarer. The point that 0% discards and 0 exposures are not the goal (§6.2.6) holds here as well. The Squad pipeline should absorb most of the bundling work, while leaving the writer time to shape the parts that matter by hand — like Doren's final line.


6.3.8 The Persona Pool — When the Same Character Appears in Multiple Cities

Once Squads stabilize, the Persona pool follows naturally as an operation. The same Persona can appear in multiple cities. An NPC from the scholars' guild being run into in three or four cities is, if anything, natural. The world does not look smaller; it looks connected.

persona_pool:
  - id: persona_doren_vale
    voice_traits: [terse, numeric]
    appearance_count: 3
    appearance_cities: [city_021, city_018, city_023]
    signature: false
  - id: persona_mira_kost
    voice_traits: [scholarly, withdrawn]
    appearance_count: 2
    signature: false

Reuse rates have a healthy range.

Reuse rate State
Under 20% NPC volume explodes; identification burden
30–50% healthy operating range
70% or more NPC staleness; diversity damage

That said, put one NPC in too many cities and players go "oh, this person again." Cap a Persona's appearances at 5 cities. Boss rooms and signature characters are no-reuse (signature: true). From a Persona's second appearance on, visual variation (lighting, props) is mandatory. This 30–50% range is not a precise number but an operating guideline — adjust it for your team and game scale.

[Signpost — If You Saw the Persona Pool as a Distribution (Still Premature)] For a team whose pool has grown to hundreds of NPCs, there is one step further — a signpost in the same spirit as the "dimension vector" section in §8.2.7 (not a prescription; for the conceptual intuition, see Appendix M). The voice_lint in §6.3.4 already measures the "closeness" of two Personas via cosine similarity (values like 0.83). Put the same embeddings over the whole pool, and staleness can be diagnosed as distribution density rather than a writer's impression — "scholarly voices piled up in one corner" becomes visible as the point density of that region. That opens a path: instead of stamping yet another similar scholar into a low-density region, you fill the gap with a variation interpolated between two nearby Personas, patching diversity that way. Two cautions, in the same breath. A Persona produced by interpolation easily becomes a "dead midpoint," an awkward blend of two NPCs — in the end a human has to bring the voice back to life. And the staleness range above (30–50%) is an operating guideline, not a precise figure; the moment you convert it into embedding distance, looseness risks passing itself off as precision — the distance value is a signal that aids the writer's judgment, not the judgment itself.


6.3.9 Six Common Failures

Failure pattern Why it fails Remedy
Mass-producing Personas, ignoring Squads all 50 NPCs exist, yet the hunting ground is dead adopt a Squad-skeleton rulebook (§6.3.3)
Free-form generation without role-distribution rules distribution accidents — five merchants, zero scholars enforce role_distribution (§6.3.3)
Relationship tags with no one-line description the relationship stays abstract, invisible in the game stage 3 AI relationship fill-in (§6.3.4)
No relationship-visibility check relationships defined, 0 exposed in dialogue or routes stage 5 관계 노출 lint (§6.3.4)
No movement-collision check 2 NPCs in the same spot at the same time, frequent after launch automatic coordinate/time-slot check (§6.3.4)
Reuse rate of 0% or 70%+ 0% explodes production volume; 70%+ goes stale pool operation + appearance cap (§6.3.8)

The fourth is the one most often missed. Defining a relationship and making it visible in the game are different jobs, and without a check the latter is almost always dropped. If the lint had not shoved relationship visibility 0 in our faces as a FAIL for silvermark hg_3, Doren and Mira would have been mentor and student only on the data sheet.


6.3.10 Try It Yourself — One Step You Can Take Today

If you're solo, just this much: You don't need a rulebook or a lint. Pick 3–5 NPCs in one location of your game (or a game you love), and write down their roles and 2 relationships by hand, in the §6.3.2 format. Then paste the relationship prompt from §6.3.4 as-is to get a one-line description, and finally ask yourself: "Where in the game's dialogue is this relationship visible right now?" If the answer is nowhere, that is exactly the lint's 관계 노출 0 (relationship visibility 0) FAIL. Close that FAIL by hand — slip a one-line mention of another member into one NPC's dialogue — and what Squad verification actually catches sinks in for good.

If you're on a team, start with this one step. Build one Squad metadata yaml form and, from the stage 5 lint, just the one-line 관계 노출 check (grep each NPC's dialogue text for other members' names and roles). The role-distribution check and the movement-collision check come after. Even with only the relationship-visibility check, you block the most common failure first: a mass-produced hunting ground shipping as a mannequin museum.

To sum up as setup → prompt → verify — setup: define roles and relationships in the Squad metadata yaml and lay the skeleton with templates.py. prompt: request the one-line relationship in the §6.3.4 format, enforcing no voice_traits clashes and no stock phrases. verify: run the stage 5 lint, confirm the 관계 노출 FAIL, and close it yourself by planting one core relationship in dialogue.


Key Takeaways

Next Chapter Preview

6.4 Content Production Workflow — Tying Multiple Generators into One Production Line

The week all three tools were finished, I sat down in one place and ran the city generator, the NPC generator, and the item generator. Each of the three worked fine on its own. I got 7 cities, 110 NPCs, and 60 weapons. It was only a few days later, sitting down to review, that I realized I had stepped into the same trap three times.

The city port_harman had been generated as a "fallen fishing village," yet the NPCs placed in it carried personas like "wealthy merchant of a thriving trade port." The city generator and the NPC generator had each looked at different lore_seeds. The item generator had stocked the shops with a level-40 legendary weapon in a city whose recommended level range is 12–18. Each of the three tools was right on its own — and wrong because they were not tied together.

This chapter is not the story of building one tool. It is an operations story: tying the city generator from 6.2, the NPC Squad from 6.3, and the item generator into a single production line. With three tools you do not get three traps; new ones grow in the gaps between the tools.


6.4.1 The Production-Line View

Run the city, NPC, and item generators separately, and each tool's output drifts out of step with the next tool's input. The fix is not making the tools smarter. It is pinning shared metadata once, upstream, and lining the tools up beneath it. If the city generator from 6.2 (chapter 2 of this part) is the model of a single tool, this chapter is the work of demoting that tool to one station on a line.

The whole line flows like this.

flowchart TD
    SEED[Mon: lore_seeds sheet
region·climate·faction·level_range]:::human SEED --> CITY[Tue: city generator
L1 rulebook skeleton + L2 AI body] CITY -->|city_manifest.json| NPC[Tue: NPC generator
inherits city lore] CITY -->|level_range| ITEM[Tue: item generator
inherits level range·faction] NPC -->|persona_pool| SQUAD[Tue: Squad placement] ITEM --> SQUAD SQUAD --> LINT[Wed: integrated lint
cross-generator consistency] LINT -->|alert| REVIEW1[Wed: first-pass review
the writer] REVIEW1 --> REVIEW2[Thu: second-pass sample
narrative lead] REVIEW2 --> BUILD[Fri: merge into build
+ prep next week's seeds] BUILD -.feedback.-> SEED classDef human fill:#fde68a,stroke:#b45309,color:#000;

The key is the arrow labeled city_manifest.json. As the city generator builds a city, it drops that city's identity — fallen fishing village or thriving trade port — into a manifest, and the NPC generator and the item generator take that manifest as input. The trap I stepped into back then existed because this arrow did not. Tying tools together means putting this one line of contract between them.


6.4.2 The One-Line Contract Between Tools — The Manifest

Every time the city generator produces a city, it also emits a city_manifest.json. Here is its actual shape. This file becomes the input for the NPC and item generators.

{
  "city_id": "port_harman",
  "display_name": "하르만 항",
  "lore_seeds": ["몰락한 어촌", "옛 무역의 잔향", "소금 부족"],
  "region": "남부 연안",
  "dominant_faction": "어민 길드",
  "level_range": [12, 18],
  "tone": "쇠락·끈질김",
  "forbidden_names": ["하란", "하르멘"],
  "neighbors": ["salt_marsh", "old_pier"]
}

The NPC generator inherits lore_seeds and tone and produces "the stubborn people of a fallen fishing village." The item generator inherits level_range and stocks only level 12–18 weapons. The forbidden_names are names already used in the city next door, so both generators avoid them. Three tools read the same single page of contract.

Here is the prompt I gave Claude to build this manifest. It is the single most important call in the production line — the one that pins down the upstream — so I quote it in full. (In this edition, the Korean data values that flow into the verbatim JSON artifacts are kept in Korean, with English glosses in parentheses.)

You are the manifest writer for an MMORPG city generator. Take the writer metadata below and generate city_manifest.json.

Writer input: - city_id: port_harman - lore_seeds: 몰락한 어촌 (fallen fishing village), 옛 무역의 잔향 (echoes of old trade), 소금 부족 (salt shortage) - region: 남부 연안 (southern coast) - level_range: 12-18

Rules: 1. Infer dominant_faction from the lore_seeds, but it must not overlap with the factions of the neighboring cities (salt_marsh, old_pier). The neighboring factions are [염전 조합 (saltworks union), 밀수단 (smuggler ring)]. 2. Put NPC/place names already used in adjacent cities into forbidden_names. Provided list: [하란 (Haran), 하르멘 (Harmen)]. 3. tone must be two words at most, and must not contradict the lore_seeds. 4. Mark inferred values separately from values the writer specified.

Claude's raw output came back like this. I am quoting it as is, not summarizing. It broke a rule in one place.

{
  "city_id": "port_harman",
  "display_name": "하르만 항",
  "lore_seeds": ["몰락한 어촌", "옛 무역의 잔향", "소금 부족"],
  "region": "남부 연안",
  "dominant_faction": "밀수단",   // ← inferred: salt shortage + old trade = smuggling
  "level_range": [12, 18],
  "tone": "쇠락·탐욕",
  "forbidden_names": ["하란", "하르멘"],
  "neighbors": ["salt_marsh", "old_pier"]
}

(Inferred values: dominant_faction, tone, display_name / writer-specified: the rest) Caution: I chose "밀수단" (smuggler ring) for dominant_faction, but under rule 1 the neighboring city old_pier is already "밀수단". This violates the rule. I suggest "어민 길드" (fishers' guild) as a replacement.

What matters here is that the AI reported its own rule violation. The smuggler ring ("밀수단") is an attractive inference from the lore_seeds — salt shortage plus old trade equals smuggling — but at the line level it is a collision. The city next door is already run by smugglers. I took the AI's suggestion: dominant_faction became the fishers' guild ("어민 길드"), and tone became "쇠락·끈질김" (decline · tenacity). "Greed" ("탐욕") was a word born from the smuggler premise and did not fit a fishers' guild.

That single verify–reject–reassign protects the upstream of the line. If the manifest is wrong, the 110 NPCs and 60 weapons below it are all generated on a false premise. Five minutes spent upstream saves three hours downstream.


6.4.3 Integrated Lint — Checking the Gaps Between Tools

A single generator's lint only looks at its own output. The city lint checks whether the city followed the rulebook; the NPC lint checks whether the personas kept voice consistency. But the trap I stepped into at the start was not inside each tool — it was between the tools. So the line needs one more layer above the per-tool lints: an integrated lint that reads cities, NPCs, and items together and cross-checks them.

Here is what the integrated lint actually catches.

Check What it compares What I had missed
Lore consistency city.lore_seeds ↔ npc.persona a wealthy merchant in a fishing village
Level-range consistency city.level_range ↔ item.required_level a level-40 weapon in a 12–18 city
Faction collision city.faction ↔ neighbor.faction two adjacent smuggler-ring cities
Name duplication the full city·npc·item name pool forbidden_names never collected

Here is part of the actual output from a run of this integrated lint. It does not auto-discard anything. It only raises alerts so that a human makes the call.

[integrated lint] port_harman line check — 3 alerts

ALERT-1 (lore consistency) port_harman
  city.lore_seeds = ["몰락한 어촌", ...]
  npc[merchant_04].persona = "번성하는 무역항의 부유한 상인"
  → Possible contradiction. Check whether this is an intended variation.

ALERT-2 (level-range consistency) port_harman
  city.level_range = [12,18]
  item[blade_legend_07].required_level = 40
  → Exceeds the recommended level range by 28. Re-review shop placement.

ALERT-3 (name duplication) — info
  npc[fisher_02].name = "하란"
  city.forbidden_names = ["하란", ...]
  → Collides with forbidden_names. Regenerating the NPC name is recommended.

ALERT-1 made me pause for a moment. An NPC who is a "wealthy merchant of a thriving trade port" is not automatically wrong. If the city used to thrive and has since fallen, then "a once-wealthy, now-poor merchant" is actually a good story. So I judged ALERT-1 not as a discard but as an "intended variation," and requested a one-line revision of the NPC persona to "an old merchant clinging to the traces of a once-thriving trade port." ALERT-2 was an outright accident, so I removed the weapon. ALERT-3 only needed the name regenerated.

The automated lint did not prevent the accident. The automated lint dragged the accident in front of human eyes, and a human made the call. That is the core of L2 (rulebook + AI assist) from 6.1. The AI builds the skeleton and the alerts; the human makes the final call. Telling apart "looks wrong but is good narrative," like ALERT-1 — a rulebook cannot do that.


6.4.4 Running the Line on a One-Week Cycle

Once the tools are tied together, the line needs a rhythm. A one-week cycle turned out to be the most stable. A week is short enough that review does not pile up, and long enough that the feedback loop is not sluggish. It lines up with one box of my desk calendar.

Day Line station Writer time
Mon Write the lore_seeds sheet (manifest upstream) half a day (5–7 cities × 15–20 min)
Tue Run the city→NPC→item generator chain no writer involvement
Wed Integrated lint + first-pass review (self) 1 hour (5–10 min per city)
Thu Second-pass sample review (narrative lead) 2–3 min per city
Fri Merge into the build + prep next week's seeds brief

Tuesday is the heart of the line. The city generator drops a manifest, the NPC generator picks it up, the item generator picks it up, and Squad handles the placement. This chain runs in the background with no writer involvement. Meanwhile the writer writes main quests (L0, fully handcrafted). This is the real payoff of tying the tools together. With separate tools, the writer has to touch the line three times on Tuesday; tied together, not even once.

One writer mass-produces 5–7 cities a week, along with their NPCs and weapons. Four weeks gives 20–28 cities. The 30-city target was reached in six weeks.


6.4.5 Line Health — Four Metrics Checked Every Week

Whether the line is healthy is read in numbers, not impressions. Four metrics, tallied automatically every week.

Metric Normal range Signal when it drifts
Integrated lint pass rate 80–95% below 60% means the manifest upstream is broken
Cross-generator collisions 3–5 per city 10+ means the contract between generators is broken
Human review discard rate 10–20% 30%+ means the production parameters are wrong
One-writer cycle time 5 days 7+ days means cognitive overload

The most line-like metric is the second one: the cross-generator collision count. When you run a single tool, this number does not exist. When it suddenly passes 10, it is not that one tool broke — it is that the contract between the tools (the manifest) broke. The usual case: the city generator's manifest schema was changed while the NPC generator was still reading the old schema. Without this metric, you do not see that accident until launch.

The four metrics are fed weekly into the quarterly retrospective. If the trend worsens, I cut the next week's city count from 5–7 down to 3–5 and look for the cause.


6.4.6 Three Accidents That Break the Line

Tie several tools together and you get accidents that a single tool never had. Here are the three I saw most often.

First, the contract-mismatch accident. A new field is added to the city generator's manifest, and the NPC generator does not know that field. The cross-collision metric spikes. When tools are developed separately, it is easy for only one side to get updated. The response is to put a version field in the manifest schema and have downstream generators raise an immediate alert on a version mismatch. You do not lecture people; you enforce the contract.

Second, the upstream-contamination accident. If the manifest is generated on a false premise (like the smuggler ring in §6.4.2), everything below it is contaminated. The human review discard rate passes 30%, yet when you look at the discarded output, the per-NPC quality is fine. The pieces are fine; the premise is wrong. The response is to add one more review at the manifest stage. Review one thing upstream rather than 110 things downstream.

Third, the model-drift accident. The LLM gets auto-updated and its output characteristics change. The city, NPC, and item generators all wobble at once. Check the past week's changes, analyze five discarded samples, and adjust the prompts or the context. Monitor for a week and confirm recovery.

The common response to all three is the same: do not blame the person; reinforce the contract. If the cause is that the writer wrote only one line of lore_seeds, then instead of saying "please write three lines," add a mandatory check to the manifest lint. That does not mean human responsibility is zero. Separately from reinforcing the system, accident patterns are shared in the retrospective.


6.4.7 Where the Writer's Time Goes

The real purpose of tying the line together is not to remove the writer, but to let the writer concentrate on signature content. One picture shows how writer time scatters when the tools are separate, and how it gathers once they are tied.

Before (separate tools) After (integrated line) Main quests 30% Signature 20% Mass-produced side review 30% Mass NPC 15% Ops 5% Main quests 50% Signature 30% Review 15% NPC 5% Ops 0% — absorbed by the line Main+signature 50% → Main+signature 80%

Main quests and signature content gather 80% of the writer's time. But this allocation does not maintain itself. Once a line is in place, writer time tends to drain into review. So I measure the time split every month, and if main-quest time falls below 50%, I reduce the number of cities produced and recover the main-quest time. The time split has to be defended as policy.


6.4.8 Extending the Line to Other Content

Once the city–NPC–item line is stable, the same skeleton extends to dungeons, collection codexes, and live events. The key is not inventing a new pattern. The temptation always comes: "dungeons are different from cities, so they need a different structure." But the skeleton of the line — shared manifest → generator chain → integrated lint → human review — is exactly the same. Only the input metadata template and the domain rulebook get swapped out.

For dungeons, a dungeon_manifest.json gains fields like boss_pattern and encounter_flow, and a domain rule such as boss routing adds one more line to the integrated lint. Same skeleton, different rules. Keeping the skeleton means the writer never has to learn yet another tool, and the integrated lint infrastructure is reused as is. That said, this is not a license to ignore domain specifics. Dungeons genuinely need routing rules that cities do not.


6.4.9 Results from Six Months of Operation

Here are the results of running this integrated line for six months on my project, compared with the period when the city, NPC, and item generators ran separately. The absolute figures below are the author's estimate (unverified), not exact tallies; the direction and the ratios follow measured trends.

Metric Tools separate After line integration
Cities produced (6 weeks) 18 28
Writer interventions on Tuesday 3 per city 0
Cross-generator collisions (found post-launch) 8–12 per quarter 2–4 per quarter
Main quests per writer per quarter 3 8
Upstream review time / downstream review time 0 / 3 hours 5 min / 1 hour

The most important change is the last row. With the tools separate, upstream review was zero and downstream review was three hours. With the line tied together and the manifest reviewed upstream, five upstream minutes erased two downstream hours. And because accidents no longer leak through the gaps between tools, post-launch consistency accidents fell from 8–12 per quarter to 2–4.

The trade-off also became explicit. Before, the abstract argument "mass production is dangerous" came around every quarter. Now decisions are made on a concrete comparison: "collisions −8 / main quests +5."


6.4.10 Seven Common Failures

1) Running the tools separately instead of tying them. The traps grow not inside the tools but between them.

2) Connecting generators without a manifest. Without a shared contract, downstream drifts away from upstream.

3) Substituting per-tool lint for the integrated lint. Per-tool lint cannot see cross collisions.

4) Compressing the cycle from five days to three. Five days is the safety margin for review.

5) Reviewing downstream while skipping the upstream (the manifest). Look at one upstream item rather than 110 downstream.

6) Charging accidents to human fault alone. Contract reinforcement and rule automation are the answer.

7) Building the line and then not using it. Enforcing the one-week cycle matters as much as the tools do.


Try It Yourself

setup. Prepare the city generator (6.2) and the NPC generator (6.3). Decide on a single city_manifest.json schema the two will share. At minimum, the fields are lore_seeds·region·faction·level_range·forbidden_names·tone·version.

prompt. Use the manifest-writer prompt from the body above as is. The last two rules are the point: "do not overlap the neighboring city's faction" (cross-collision prevention) and "mark inferred values separately from specified values" (reviewability). When a city is generated, have it drop a manifest, and wire the NPC and item generators to take that manifest as their input.

verify. Run the integrated lint once. It cross-checks four things: lore consistency, level-range consistency, faction collision, and name duplication. When an alert fires, do not auto-discard — have a human judge it. Telling apart "looks wrong but is good narrative" (the old merchant of a fallen trade port) is the human's share of the work.

Solo Scale-Down. Even if your only tools are a city generator and an NPC generator, the line still holds. Put each city's lore_seeds, level_range, and forbidden_names on one row of a spreadsheet, and pasting that whole row into the NPC generator prompt already plays the manifest's role. An integrated lint with a single rule — city–NPC lore consistency — is enough to block that trap. No grand infrastructure required: one principle, "put one line of contract between the tools," is enough to start the line.


Key Takeaways

Part 7 · Level Design

7.1 Procedural Level Design Master

The exit of room 47 in the dungeon was sealed. The build had passed. QA had passed. It was the third day of live service when a screenshot appeared on the community boards: a player standing in front of a wall, one room short of the boss. That room had been copy-pasted from a hand-built room two quarters earlier, and somewhere in the copying, one eastern corridor kept its visuals but lost its connection data. Nobody verified it. There was no tool to verify it with.

This chapter is the story of building a structure that blocks that accident automatically, at the build stage. The heart of it is not the craft of drawing spaces — it is the practice of running the data attached to those spaces as rules.


A level design workshop is closer to a drafting room. Each blueprint comes from someone's hands, but consistency, reuse, and verification across blueprints are decided by how the blueprint cabinet is run. Anyone can build one hand-crafted dungeon. Running a hundred dungeons with a consistent difficulty curve and a graph free of dead ends is not a matter of craft — it is a systems problem.

On Project A — the MMORPG where I work as design director (targeting Korea and Southeast Asia, a mid-size team of 10–50, mobile-first) — this system has a name: a single document called Procedural_Level_Design_Master. This chapter covers what that document unifies, how far AI gets to touch it, and where AI stops. Underneath this chapter sits my experience leading design on a mobile roguelite RPG where the dungeon was generated fresh every run — running procedural space as rules.

7.1.1 Two Paths — Generate the Space, or Operate Its Metadata?

Level automation splits in two directions. One is generating the space itself procedurally. Classic PCG (Procedural Content Generation) lives here: BSP partitioning (Binary Space Partitioning, a classic technique that recursively bisects a space to place rooms), wave function collapse, drunkard's-walk grids. The other is operating the space's metadata — room tags, connectivity, difficulty labels, event slots.

Classic PCG is strong at the first. In genres where "a new map every run" is the core of the game — roguelikes, sandboxes — the first is the right answer. MMORPGs are different. Players run the same dungeon dozens of times. They run it until the route is memorized. So the dungeon needs to be a fixed, hand-polished space, and the place for automation is not the space itself but the metadata that makes that space operable.

Why metadata is the spine of operations becomes clear deliverable by deliverable.

Deliverable Without metadata
A pool of dozens of dungeons No way to search which room is where; no reuse
Difficulty curve validation No per-room difficulty labels, so no curve can be drawn
Automatic quest and boss placement No event slot metadata, so coordinates are entered by hand
Art team sync No room type → art set mapping, so visuals drift apart
Measuring player routes and dwell time No room-ID-based telemetry

A dungeon without metadata builds, but it cannot be operated. It is a library full of books with no index. That is why this chapter focuses on "operating spatial metadata."

7.1.2 What the Master Document Unifies

Procedural_Level_Design_Master binds four standards into one document: the room metadata format, the room tag dictionary, the connectivity rules, and the validation checklist. Start with what happens when these four are scattered. When five designers each reference the format from a different file, one writes the type field as combat, another as Combat, another as battle_room. Search breaks, statistics break, and eventually automation breaks.

Sorted by Layer, each of the four standards has an obvious home. The format, the dictionary, and the rules belong to the rulebook that governs generation (L1); the generated room bodies are content (L2); sheet values are data (L3); validation sits at the build/QA gate (L4).

L0 Vision Level concept · pacing intent (immutable anchor, injected into every generation/validation) — art_pack tone, difficulty intent L1 System Rulebook — room meta format · tag dictionary · connectivity rules — the slot the Master document binds L2 Content Room bodies with metadata attached (generated · polished spaces) L3 Data Room size · connection sheets · event slot IDs · enemy data L4 Build · QA Graph validation · difficulty curve validation · art set consistency gate

When I say the master document unifies the four standards, I do not mean "cram all the body text into one file." I mean "collect the rules in the L1 slot." That is what lets the automation that comes later sit on top of the Layer boundaries (what happens when the separation collapses is covered in 7.1.11).

7.1.3 The Room Metadata Format — the Input Slot Where Automation Attaches

Each room follows the format below. This format is the input interface for automation.

room_id: dungeon_021_room_07
dungeon: dungeon_021_silvermark_library
type: combat_room          # combat / puzzle / lore / safe / boss
size: medium               # small / medium / large
difficulty_label: hard_for_level_28
tags: [scholar_theme, vertical_layout, water_hazard]
connections:
  - target_room: dungeon_021_room_06
    type: door
    direction: south
  - target_room: dungeon_021_room_08
    type: passage
    direction: east
event_slots:
  - slot: enemy_spawn_1
    constraints: [scholar_enemy, level_28]
  - slot: lore_object_1
    constraints: [scholar_lore]
movement_complexity: 4     # 1~5
estimated_clear_time_sec: 90
art_pack: scholar_library_v2

Every field has at least one automation consumer. type feeds dungeon pool statistics and difficulty calculation; tags feeds search, reuse, and art set mapping; connections feeds graph validation (dead-end checks); event_slots feeds automatic quest and boss placement. A field with no consumer does not go into the format. It only adds input cost without adding value.

7.1.4 The Room Tag Dictionary — Small and Orthogonal

Tags are the metadata's search keys. Let them multiply without limit and search breaks. Put 200 labels on a cabinet of drawers and you can no longer find what is where. So we run 5 categories × roughly 6 enums per category — about 30 in total.

Category Enums Examples
theme 8 scholar_theme, ruins_theme, forest_theme …
layout 5 vertical_layout, horizontal_corridor, open_arena …
hazard 6 water_hazard, fire_hazard, falling_hazard …
interaction 4 puzzle_required, lever_activation …
narrative 7 flashback_trigger, dialogue_zone …

A room never carries more than 5 tags; 3–4 is normal. Adding a new tag means passing a four-step gate: it must be a candidate for use in 5 or more rooms per quarter, it must be inexpressible as a combination of existing tags, its use in search or art set mapping must be clear, and it must still be on 5 or more rooms after a month in operation. The last condition is the key one. A tag created on impulse, used once, and abandoned pollutes the dictionary.

7.1.5 The Procedural Level Pipeline — From Rulebook to Validation

How the standards so far connect into one flow — that connecting line is the skeleton this chapter rests on. It is a pipeline that starts at the rulebook, passes through AI-assisted variation, and ends at guardrail validation.

flowchart TD
    A["L0 vision — dungeon concept · pacing intent"] --> B["L1 rulebook\ntag dictionary · connectivity rules · slot rules"]
    B --> C["Room skeleton layout\ndesigner handwork + editor"]
    C --> D["Automatic metadata extraction\nroom_id · connections · type · size"]
    D --> E["AI-assisted variation\ntag extraction · art_pack mapping suggestions"]
    E --> F{"Dictionary enforcement check\ntags outside the dictionary?"}
    F -->|outside dictionary| E
    F -->|pass| G["Designer review\nfinalize tags · difficulty_label"]
    G --> H["L4 graph validation\nreachability · dead ends · loops · branching"]
    H -->|violation| C
    H -->|pass| I["Difficulty curve validation\nsum of room difficulty labels"]
    I --> J["Art set consistency gate"]
    J -->|pass| K["Build — register in dungeon pool"]
    J -->|mismatch| E

    classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545;
    classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764;
    classDef human fill:#fde68a,stroke:#b45309,color:#000;
    classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d;
    class B,D,F,H,I,J code;
    class E ai;
    class C,G human;
    class K pass;

Three properties of this pipeline are worth pinning down. First, the rulebook (L1) sits upstream of all generation. Second, AI varies only within the dictionary the rulebook defines — the F gate sends out-of-dictionary output back. Third, validation (H, I, J) is fixed as a gate just before the build, so violations are blocked by code rather than depending on anyone's attention. The room 47 accident happened because there was no H gate.

7.1.6 Connectivity Rules — Guardrails Verified as a Graph

The connections field in the room metadata turns the whole dungeon into one directed graph. Once it is a graph, validation is automatic.

Check On violation
Start room → boss room reachable Blocked as a build failure
Dead end (1 exit + non-safe_room) alert — designer review
Bidirectional link consistency (A→B exists but no B→A) Auto-corrected
Loop length — short 2–3-room loops alert
Branching width — 4 or more simultaneous branches Designer review

The measurement script takes the following shape: a thin wrapper that lays dungeon vocabulary over standard graph algorithms (longest path, average out-degree, loop count, shortest path).

# level_graph_metrics.py
def measure(dungeon):
    graph = build_graph(dungeon.rooms)
    return {
        "depth":            longest_path_length(graph),
        "branching_factor": avg_out_degree(graph),
        "loop_count":       count_loops(graph),
        "dead_ends":        count_dead_ends(graph),
        "boss_reachability": shortest_path(graph.start, graph.boss),
    }

The five metrics come out in a form comparable across dungeons. We use them as diversity metrics for the dungeon pool. But diverse metrics do not mean a fun dungeon. Metrics exist to block accidents, not to guarantee fun. Zero dead ends guarantees no fun at all. Fun comes from a designer's insight; graph validation just holds up the floor so that insight does not get buried under accidents.

7.1.7 Worked Example — Handing tags Extraction to AI, Rejecting It, and Asking Again

Of all the automation, the part people most often want to take their hands off is tags entry. Tagging 100 rooms is tedious, and even a human gets confused looking at room screenshots alone. Repetitive work with clear judgment criteria is exactly the kind of place where AI is good at carrying the first draft. This section lays out the workflow I actually ran — the prompt, the AI output that got rejected, and the human's follow-up request — with nothing retouched.

First prompt:

[Input]
- Room screenshot: (attached)
- Room area: 18m × 12m, ceiling height 9m (two-story vertical structure)
- Placed enemies: scholar_phantom × 3, water_elemental × 1
- Adjacent room type: lore_room (west), combat_room (east)
- A shallow pool of water inside the room

Pick 3~5 tags that fit this room, only from the dictionary below. Don't invent tags
outside the dictionary; output the tag enums only, comma-separated.

[Tag dictionary, 30 entries]
theme: scholar_theme, ruins_theme, forest_theme, cave_theme,
       crypt_theme, swamp_theme, mechanical_theme, celestial_theme
layout: vertical_layout, horizontal_corridor, open_arena,
        maze_layout, bridge_layout
hazard: water_hazard, fire_hazard, falling_hazard,
        poison_hazard, crush_hazard, dark_hazard
interaction: puzzle_required, lever_activation, key_gate, hidden_passage
narrative: flashback_trigger, dialogue_zone, lore_object_zone,
           ambush_setup, ritual_site, vista_point, rest_anchor

Claude's first output (raw, untouched):

scholar_theme, vertical_layout, water_hazard, ambush_setup, two_story_layout

Two problems are exposed right here. One: two_story_layout is not in the dictionary. The model got pulled by the input phrase "two-story vertical structure" and invented a word outside the dictionary. Two: the tags hit the cap of five, yet vertical_layout and two_story_layout overlap in meaning — two slots spent on the layout category.

The human's rejection call. two_story_layout violates dictionary enforcement, so it is an automatic reject (the pipeline's F gate). ambush_setup is a closer call — three enemies are placed, but from the screenshot it is ambiguous whether this is an ambush composition or just a combat room. The model may have over-interpreted. With the adjacent rooms both lore/combat, the case for a narrative ambush trigger is weak. Rejected.

Second request:

two_story_layout isn't in the dictionary and overlaps with vertical_layout. Drop it.
Drop ambush_setup too — having enemies doesn't make it an ambush.
// (intent: narrative tags only when narrative intent is explicit; plain combat is already expressed by type)
You don't have to fill the empty slots. 3~4 is normal. Tag it again.

Claude's second output:

scholar_theme, vertical_layout, water_hazard

This time it is down to three, all inside the dictionary, with no category duplication. The human accepted the result. That is how the room format's tags: [scholar_theme, vertical_layout, water_hazard] got finalized.

The lesson of this worked example is two lines long. First, AI overfits to one detail of the input ("two-story") and steps outside the dictionary — the dictionary enforcement gate has to catch that at the code level. Second, AI tends to fill empty slots — unless you state "you don't have to fill them," it will try to use all five. Both failures are common, and both fixes are only stable when enforced not by the prompt but by the rulebook (dictionary + cap).

7.1.8 Metadata at Scale — Who Fills It In, Who Reviews It

Filling in one room's metadata by hand takes a designer 5–10 minutes. One dungeon (20–30 rooms) is 2–5 hours; 100 dungeons is 200–500 hours (author's estimate, unverified — an upper bound computed as average entry time per room × room count). Fill it all by hand and the designer becomes a metadata entry slave.

So the work is split by area, each with its own filler.

Area Filled by
room_id · dungeon · connections Automatic editor extraction (L3)
type · size Auto-classified from room area and connection count
tags AI assist + designer review (7.1.7)
event_slots Rulebook per room type
difficulty_label Auto-computed by summing the enemy data in the room
art_pack Room type · dungeon theme mapping

What designers finalize by hand comes down to reviewing tags and giving final sign-off on difficulty_label. Tools fill in the rest; people review. The point of the automation is to pull designers out of data entry and send them back to the judgment calls — pacing, signature rooms, reuse policy.

7.1.9 Room Reuse and Its Pitfalls

The biggest payoff of the master standard is room reuse. With 30 rooms searchable by tag, you can assemble 5–10 dungeons out of combinations. But as the reuse ratio climbs, dungeons go stale. So reuse ships with guardrails attached.

Guardrail Definition
A room appears in at most 5 dungeons Appearance count tracked automatically
Visual variation required on second appearance Lighting and prop changes
No reuse of boss rooms or signature rooms Enforced by flag
Track negative feedback on reused rooms Player telemetry

Reuse is a means of cutting cost, not a goal. The moment the reuse rate itself becomes a KPI, the player experience flattens. At 0% (every room new), production cost explodes; past 70%, dungeons stop being distinguishable from one another. In my experience the 30–40% band is the balance point between cost and variety (a directional observation — the precise threshold differs by project).

7.1.10 Common Failures and Fixes

Pattern Fix
Five people read the metadata format five different ways Unify it at L1 with the Master document
Tags multiply to 50–100 The 30-tag dictionary + the four-step gate
Builds ship without dead-end checks Make graph validation a build gate
Designers hand-enter all metadata Editor extraction + AI assist
AI generates out-of-dictionary tags Auto-reject at the dictionary enforcement gate
Reuse at 0% or 70%+ The 30–40% band + variation guardrails

7.1.11 Layer Decomposition Is the Precondition for Procedural Level Generation

The structure of 7.1.2–7.1.6 — laid out so far as rulebook, generation, validation — is itself a product of Layer decomposition. The general thesis that Layer decomposition is the precondition for procedural generation and automation (L0 anchor → L1 rulebook → L2 body → L3 values → L4 gates; as one lump, generation collapses) was covered in §6.6. Here I apply it to operating level metadata.

Without this separation, room layout, BSP, pacing, and narrative triggers all mix into one file, and every time you move a single room, the pacing intent, the event slots, and the connectivity graph break together. It is the drafting room, the materials warehouse, and the inspection room piled onto one desk — pull out one blueprint and the materials invoice and the inspection sheet come out with it. The AI assist in 7.1.7 worked because of the Layers, too: room IDs and connectivity are filled from the editor (L3 auto-extraction), tags from AI (L1 dictionary enforcement), difficulty_label from summation (L3→L4). Automation sits on Layer boundaries; sit it on a single lump and accidents multiply within the first quarter and the tool itself gets thrown out.

That said, this does not mean you need the full five-drawer cabinet in perfect shape from day one. Separate gradually; keep interfaces narrow. In the first quarter, separating just the L1 rulebook (tag dictionary + connectivity rules) and the L3 sheet (room metadata sheet) already creates a place for automation to enter. The L0 pacing intent and the L4 validation gates get filled in over the quarters that follow. Once the standards are unified, automation has a place to land — and the more automation lands, the more the designer's attention shifts from hand-working one room at a time to judging pacing, signature rooms, and reuse.


Key Takeaways


Try It Yourself

setup. Pick one dungeon and build a YAML sheet where each room carries only four fields: room_id · type · connections · tags. First pin the tag dictionary — 5 categories, about 30 enums — on a single sheet of paper.

prompt. Feed in a room screenshot + area + enemy types + adjacent room types, and ask: "pick only 3–5 tags from this dictionary, no tags outside the dictionary, don't fill empty slots" (the exact prompt from 7.1.7).

verify. (1) If the AI output contains a tag outside the dictionary, reject it and ask again. (2) Build a graph from connections and check start→boss reachability and dead ends — if even one violation turns up, mark that room as unbuildable.

Solo Scale-Down

If you are a solo developer with no tooling infrastructure, start the master document as a single markdown page. Thirty lines of tag dictionary, five lines of connectivity rules, five lines of validation checklist — that is enough. For graph validation, if you have 10 rooms or fewer, drawing arrows on paper and eyeballing for dead ends gets you 80% of the effect. The core is not the tool but the habit itself: attach data to rooms, and check that data with rules. Add the tools when you pass 50 rooms and checking by hand becomes too heavy.

Next Chapter Preview

7.2 The Behavior Tree Editor — A Worked Transcript of a Human and AI Editing and Verifying BT json Together

An apprentice mage was glued to the player, swinging a sword. I had designed that NPC as a ranged magic caster. Its HP was paper-thin — one melee hit and it was dead — yet it had no intention of keeping its distance. The build log showed no errors. I reopened the Behavior Tree in the editor; the nodes were all wired up correctly. After an hour of staring at it, I found the cause. The distance condition on the retreat branch was 0.5 instead of 5. The NPC was supposed to flee when an enemy came within 5 meters, but at 0.5 meters — practically point-blank — the retreat branch never fired.

One number. In the graphical node editor, that number was visible only when you expanded the node's inner panel, and it left nothing in the change history. There was no way to trace who changed that value, or when. From that day on, my Project A started handling Behavior Trees as json rather than graphics. This chapter is the record of one cycle in which a human and AI edit that json together and a machine verifies it automatically.


7.2.1 Where the Behavior Tree (BT) Slips Out of Your Hands

The Behavior Tree is the de facto standard structure for defining an enemy NPC's combat, movement, and reactions. A selector tries branches in priority order; a sequence chains conditions and actions in order. The structure itself is simple. The problem is scale.

On my Project A, a single enemy NPC's BT ran roughly 50–200 nodes, and we operated more than 100 NPCs. Multiply those out and the total BT node count reaches tens of thousands. At that scale there comes a moment when no human can answer the question, "If I change this retreat pattern, which NPCs are affected?" It is like having a hundred notebooks spread open on your desk: fix one line in the first volume, then try to trace by eye where it bleeds into the other ninety-nine.

When I moved from graphical BTs to json, I demanded four things.

Store as text (json) Track every changed line via git diff "One number" incidents stay in history Standardized node metadata Search and reuse via category·tags "Find similar BTs" becomes a one-line query subtree references (reuse by reference) One common pattern shared by many BTs Fix one place, not copy-paste → applies everywhere Automatic change-impact visibility Which BTs a subtree edit reaches A script computes it, not human guesswork

The built-in BT editors in commercial game engines integrate easily and offer strong visual debugging. They do, however, tend to store BTs as binary assets, which weakens text diffs and change-impact tracking. Project A assumed a live game operating more than 100 BTs, so we chose to develop our own json BT format and editor. Let me be clear: this is not the right answer for every team. If you operate fewer than 50 BTs, sticking with the engine's built-in editor is almost always cheaper. I come back to the justification for in-house development at the end of this chapter.


7.2.2 BT json — One Enemy's Behavior as Text

Start with what the result looks like. Below is part of the BT for a ranged-support NPC of the Scholars' Guild. Two things are key: every behavior is text, so git can track it line by line, and common patterns are referenced via subtree_ref.

{
  "bt_id": "bt_scholar_archer_v3",
  "category": "ranged_combatant",
  "tags": ["scholar_faction", "ranged", "support"],
  "description": "Scholars' Guild ranged support type. Keep distance + retreat first.",
  "root": {
    "type": "selector",
    "children": [
      {
        "type": "sequence",
        "name": "low_hp_retreat",
        "children": [
          {"type": "condition", "fn": "hp_below", "param": 0.3},
          {"type": "subtree_ref", "id": "subtree_retreat_to_ally"}
        ]
      },
      {
        "type": "sequence",
        "name": "kite_pattern",
        "children": [
          {"type": "condition", "fn": "enemy_in_close_range", "param": 5},
          {"type": "action", "fn": "move_away", "param": {"distance": 8}}
        ]
      },
      {"type": "subtree_ref", "id": "subtree_ranged_attack_pattern"}
    ]
  }
}

Unfolded as a diagram, the tree is a selector trying three branches from the top. Note that the bug from the opening — whether the param of enemy_in_close_range is 5 or 0.5 — becomes a single line you can see at a glance in json.

flowchart TD
    R["selector
(priority from the top)"] R --> A["sequence: low_hp_retreat"] R --> B["sequence: kite_pattern"] R --> C["subtree_ref:
subtree_ranged_attack_pattern"] A --> A1["condition: hp_below 0.3"] A --> A2["subtree_ref:
subtree_retreat_to_ally"] B --> B1["condition: enemy_in_close_range 5"] B --> B2["action: move_away dist=8"] style C fill:#e8f0fe,stroke:#4285f4 style A2 fill:#e8f0fe,stroke:#4285f4 style B1 fill:#fce8e6,stroke:#ea4335
Element Role
bt_id Key for git diff and change tracking
category / tags Unit of search and reuse
subtree_ref Reference to a common pattern (fix one place → update many BTs)
description Shared with designers and scenario writers

The node painted red, enemy_in_close_range 5, is the one that ate an hour of a human's time in the opening. In json, a single code review catches it.


7.2.3 The subtree Library — Reference Instead of Copy-Paste

Across the behaviors of more than 100 enemies, recurring chunks appear: "retreat behind an ally," "retreat to cover," "ranged attack pattern," and the like. Copy-paste those into every BT, and fixing one piece of retreat logic means hand-hunting a hundred places. So common patterns are split off into separate subtree files and only referenced via subtree_ref.

subtree_library/
├── retreat_patterns/
│   ├── subtree_retreat_to_ally.json
│   ├── subtree_retreat_to_cover.json
│   └── subtree_retreat_random.json
├── attack_patterns/
│   ├── subtree_ranged_attack_pattern.json
│   ├── subtree_melee_combo.json
│   └── subtree_aoe_attack.json
└── reaction_patterns/
    ├── subtree_react_to_ally_death.json
    └── subtree_react_to_player_taunt.json

Set up this way, "who is affected if I change this subtree?" becomes the output of a script, not a human's guess. The impact tracker is simple: open every BT and collect the bt_id of each BT that references the subtree in question.

# bt_impact_tracker.py
import json, glob

def has_subtree_ref(node, target_id):
    if isinstance(node, dict):
        if node.get("type") == "subtree_ref" and node.get("id") == target_id:
            return True
        for child in node.get("children", []):
            if has_subtree_ref(child, target_id):
                return True
    return False

def find_affected_bts(subtree_id):
    affected = []
    for bt_file in glob.glob("bts/*.json"):
        bt = json.load(open(bt_file, encoding="utf-8"))
        if has_subtree_ref(bt["root"], subtree_id):
            affected.append(bt["bt_id"])
    return affected

# usage
affected = find_affected_bts("subtree_ranged_attack_pattern")
# → ["bt_scholar_archer_v3", "bt_ranger_v2", "bt_sniper_v1", ...]

On Project A we wired this function into the pull request (PR) stage. When someone touches a subtree file, the list of affected BTs is automatically posted as a PR comment. Before the merge, the reviewer sees that "one changed line in a retreat pattern alters all 12 ranged enemies."


7.2.4 Worked Transcript — One Cycle of AI Drafting a New BT

This is where the chapter carries the most weight. I hand the BT draft for a new enemy NPC, an "apprentice mage," to AI, then reproduce — unpolished, exactly as it happened — one cycle of a human verifying, rejecting, and re-requesting that output. There is a reason I do not smooth it over: what the AI got wrong on its first output, and the texture of that failure, is everything this chapter is trying to convey.

Step 1 — The Human's Prompt (Full Text)

[Enemy NPC concept]
Apprentice mage of the Scholars' Guild. Low HP, high magic damage.
Avoiding melee comes first. Retreat behind an ally, then attack with magic.

[category] ranged_combatant

[Reference BT] bt_scholar_archer_v3
(Similar structure. But replace the bow attack with a magic attack)

[Defined function enum]
condition: hp_below, enemy_in_close_range, ally_nearby, mana_below
action: move_away, cast_spell, retreat, wait
subtree_ref id (available list):
- subtree_retreat_to_ally
- subtree_ranged_attack_pattern   # Bow-based. Not for magic
- subtree_magic_attack_pattern

Write only the BT root node json for this concept. Use subtree_ref as much as possible, keep direct nodes minimal,
functions only from the enum above. Do not add behaviors not in the concept. json only, no explanation.

Step 2 — Claude's First Output (Verbatim, Unedited)

{
  "bt_id": "bt_scholar_apprentice_v1",
  "category": "ranged_combatant",
  "tags": ["scholar_faction", "magic", "fragile"],
  "description": "Scholars' Guild apprentice mage. Avoid melee + retreat behind ally + magic attack.",
  "root": {
    "type": "selector",
    "children": [
      {
        "type": "sequence",
        "name": "low_hp_retreat",
        "children": [
          {"type": "condition", "fn": "hp_below", "param": 0.4},
          {"type": "subtree_ref", "id": "subtree_retreat_to_ally"}
        ]
      },
      {
        "type": "sequence",
        "name": "avoid_melee",
        "children": [
          {"type": "condition", "fn": "enemy_in_close_range", "param": 6},
          {"type": "action", "fn": "move_away", "param": {"distance": 10}}
        ]
      },
      {
        "type": "sequence",
        "name": "mana_recover",
        "children": [
          {"type": "condition", "fn": "mana_below", "param": 0.2},
          {"type": "action", "fn": "wait", "param": {"frames": 120}}
        ]
      },
      {"type": "subtree_ref", "id": "subtree_ranged_attack_pattern"}
    ]
  }
}

Step 3 — Human Verification and Rejection

I ran the output through lint (covered in the next section) and read it with my own eyes. Three things snagged.

  1. The last branch is wrong. subtree_ranged_attack_pattern is a bow-based attack pattern. Even though the enum comment said "not for magic," the AI pulled in the bow pattern with the similar name. The apprentice mage should use subtree_magic_attack_pattern. — Rejected.

  2. It added behavior not in the concept. The mana_recover branch (mana_belowwait 120frame) is a behavior with not a single line in the concept. The prompt nailed down "do not add behaviors not specified," yet the AI apparently judged that "a mage would naturally manage mana" and slipped it in. Plausible, but a spec violation. — Rejected.

  3. The retreat threshold is suspect. hp_below 0.4 starts retreating earlier than the reference BT (0.3). The apprentice mage is indeed more fragile, so it could be intentional. But the AI never gave a rationale, so this is the designer's call. — On hold, pending designer confirmation.

Step 4 — The Human's Follow-Up Request

Fix just three places. The last branch, subtree_ranged_attack_pattern, is the bow pattern,
so swap it for subtree_magic_attack_pattern; the mana_recover branch is not in the concept, so delete it entirely;
and revert hp_below 0.4 to the reference BT (0.3) — if there is a rationale for 0.4, put it in the description in one line.
Keep the rest as-is, json only.

Step 5 — Claude's Second Output (Revised)

{
  "bt_id": "bt_scholar_apprentice_v1",
  "category": "ranged_combatant",
  "tags": ["scholar_faction", "magic", "fragile"],
  "description": "Scholars' Guild apprentice mage. Avoid melee + retreat behind ally + magic attack. Low HP, so retreat threshold kept at 0.3 (same as archer).",
  "root": {
    "type": "selector",
    "children": [
      {
        "type": "sequence",
        "name": "low_hp_retreat",
        "children": [
          {"type": "condition", "fn": "hp_below", "param": 0.3},
          {"type": "subtree_ref", "id": "subtree_retreat_to_ally"}
        ]
      },
      {
        "type": "sequence",
        "name": "avoid_melee",
        "children": [
          {"type": "condition", "fn": "enemy_in_close_range", "param": 6},
          {"type": "action", "fn": "move_away", "param": {"distance": 10}}
        ]
      },
      {"type": "subtree_ref", "id": "subtree_magic_attack_pattern"}
    ]
  }
}

The second output passed lint; in designer review we nudged the avoid_melee distance 6 down a notch to 5 to match the archer, then sent it on to the simulation gate. The lesson here is clear. AI fills in a BT's skeleton fast, but it commits two accidents almost every time: pulling in the wrong subtree with a similar name, and slipping in plausible extra behavior with no spec behind it. Only human eyes and the lint gate catch these two. That is why AI output is a draft, not a final.


7.2.5 Automated lint — The Machine Catches What Humans Miss, First

BTs connect directly to the player experience. Ship the bug where an enemy will not flee at point-blank range, and it comes back as review scores. So before the merge, the machine checks first.

Check On violation
Unreachable node alert (a branch the selector can never reach)
Infinite loop risk block (a repeating sequence with no exit condition)
subtree_ref target missing block
Action/condition function outside the enum block
Node count explosion (>500) alert (recommend splitting the BT)
Response-time variance across BTs in the same category alert (suspected balance regression)

The last row is what makes this lint unusual. If the five BTs grouped under the same ranged_combatant drift far apart in average simulated response time, that is a signal that someone silently broke one enemy's balance. It is a device that catches, with statistics, the "vibe" that static checks cannot.

After static lint comes simulation verification. Run the BT 1,000 times in a simulator — no actual game build required — and pull the statistics.

Metric Normal range
Average survival time (vs. a standard player) Per-category baseline
Attack pattern diversity (entropy) 0.6 or higher
Retreat/approach behavior ratio Per-category baseline
Average frames per action 60 frames or fewer

Without baking a build, you can see within 5–10 minutes whether this BT dies too fast or keeps repeating a single behavior. If an anomaly shows, fix the json and re-run the sim. The real gain of going json is that this cycle shrinks from days to minutes.

flowchart LR
    P["Human/AI
edits BT json"] --> L{"Static lint"} L -->|block| P L -->|pass| R["Designer review"] R -->|reject| P R -->|approve| S{"1,000 sim runs"} S -->|anomaly| P S -->|normal| M["Merge + build"] style L fill:#fef7e0,stroke:#fbbc04 style S fill:#fef7e0,stroke:#fbbc04 style M fill:#e6f4ea,stroke:#34a853

7.2.6 Measurement — What Shrank

Here is Project A before and after adoption, as a table. The absolute figures vary with team size and game genre, so they are the author's estimates (unverified). The direction and the ratios, however, are exactly what we observed in live operation.

Item Before (engine built-in BT, direct) After (json + editor)
Writing a new enemy's BT 1–2 days 2–4 hours
Assessing BT change impact Guesswork and experience Automatic (subtree impact list)
Verification after a change Real build required 5–10 min simulation
Operating 100 enemy NPCs 3 designers full-time 1–2 designers
Post-launch BT incidents (abnormal behavior) 10–15 per quarter (author's estimate) 2–4 per quarter (author's estimate)

What matters most is that the last two rows moved at the same time. Normally, cut headcount and quality drops. Here, the number of designers went down and incidents went down with it — because the machine took over the change-impact tracking and verification that people had been doing by hand. The value of automation lies less in "faster" than in this "smaller and better at once."


7.2.7 Build In-House, or Borrow?

If this chapter leads you to conclude "we should build a json BT editor too," that is the wrong takeaway. Project A chose in-house development because a specific set of conditions lined up.

Option Pros / Cons
Use the engine's built-in BT as-is Easy integration / weak json conversion and diffs
Adopt an external BT library Standardization benefits / learning curve, customization limits
In-house json BT editor + runtime Best freedom and traceability / high development cost

Project A picked option 3 for four reasons.

Development cost: 1–2 months. It pays back only when operated BTs reach 100–300 and the live ops period runs long. At a scale of 30–50, it does not pay back. The return on investment (ROI) of in-house development materializes only when both scale and operating period are assured. If your team is small, take only the principles from this chapter — store as json, reference via subtrees, pass AI output through the lint-plus-review gate — and run your tooling on top of the built-in editor or an external library.


7.2.8 Common Failures

Pattern Remedy
Managing BTs only as binary assets Store them as json to keep git tracking alive
Copy-pasting the same pattern into every BT, no subtrees Split it into a subtree library and reference it
Tracking BT impact by hand Wire the impact analysis script into PRs
Verifying only in real builds, no simulation Run a simulator decoupled from the build
Using AI-output BTs without review Pass them through the triple gate: lint + designer + simulation
Never measuring in-house development ROI Build in-house only at 100+ BTs in live ops

Key Takeaways


Try It Yourself

This is the smallest cycle a small team can try today.

setup — Write out the BT of one enemy NPC you currently operate as json, by hand (bt_id, category, tags, root). Pull one common retreat or attack pattern out into subtree_library/ and reference it with subtree_ref.

prompt — Hand a similar new enemy to AI. Use the prompt skeleton from the worked transcript above as-is: concept + category + reference BT + available function enum + "no behaviors not specified" + "json only."

verify — Merge AI output only after it passes three gates: (1) a lint that filters out functions outside the enum and nonexistent subtrees, (2) human eyes, (3) a simulation or a short in-game check. Always confirm whether the AI slipped in "a wrong subtree with a similar name" or "plausible out-of-spec behavior."

Solo Scale-Down

If you have no capacity to build an editor, your tools are a text editor, git, and one 30-line bt_impact_tracker.py — that is enough. Export a BT built in the engine's editor to json once, commit it to git, and split only the subtrees into separate files to reference. Hook the impact tracking script into a commit hook, and even working alone you can see which enemies change when you touch a retreat pattern — as output, not as a guess. This one habit alone shrinks the opening's "one number, one hour" down to a single line in code review.


Next Chapter Preview

7.3 The Dungeon and Field Pattern Library

At a dungeon review, a junior level designer put one of their dungeons up on the screen. A narrow corridor, a fast enemy closing in from behind, a dodge decision at the junction. It was a well-made dungeon. The problem was that it was subtly different from what we had already built in eleven other dungeons. The enemy's pursuit speed, the timing of the trap going off, the moment the junction appears. Not one of them was the same. The junior designer believed they had built an experience with the same name — "pursuit dungeon" — but the sensation players received differed from dungeon to dungeon.

What we decided that day was simple. Define the "corridor pursuit" experience precisely, once, and pin that definition down. The next time someone builds a pursuit dungeon, they don't compose it from scratch — they pull out the codified definition and use it. That was the beginning of the pattern library.

If the room is the unit of space and the Behavior Tree is the unit of behavior, the pattern is the operational unit that binds space, behavior, and events into one. When a single pattern is reused across several dungeons, the production load drops — and more importantly, the experience players receive stays consistent across dungeons.


7.3.1 The Pattern as an Operational Unit

Think of a recipe in a cookbook — the analogy is exact. One recipe card holds the ingredients, the cooking steps, the heat level, and a photo of the finished dish. Even when the restaurant changes, following the same recipe produces the same taste. Each restaurant is allowed a little variation, though. A pattern is the same. Space (the rooms), behavior (the BT subtrees), incidents (the events), outcomes (reward and difficulty), and the designer's statement of intent go in as one bundle.

Pattern = a bundle of five elements Space Room metas 1~3 Behavior BT subtrees 1~2 Events event slots Outcome Reward · difficulty rules Intent Description "Corridor pursuit pattern" = narrow corridor + fast-enemy BT + trap event + dodge reward → Once verified, reproduces the same experience across 5~10 dungeons

Once a pattern is defined, the same experience can be built consistently in every dungeon — the way a recipe proven once delivers the same taste across many restaurants. Still, even with the same recipe, each restaurant keeps a little variation. How that variation is managed is half of pattern operations. The overrides covered later are where it lives.


7.3.2 The Pattern Composition Flow

The heart of a pattern library is this: patterns are codified into a rulebook first, and dungeons are then generated by combining them. The designer does not compose a dungeon on a blank screen; they pick verified patterns, place them, and vary only a part.

flowchart TD
    A[Observe a good experience moment in the game] --> B[Decompose into space · NPCs · events]
    B --> C[Map to room templates · subtrees]
    C --> D{Pass simulation + user testing?}
    D -- No --> B
    D -- Yes --> E[Register the pattern in the library
usage_count = 0] E --> F[(Pattern library
5 categories · 30–50 patterns)] F --> G[Dungeon design: invoke pattern instances] G --> H[Placement + variation overrides] H --> I{Variation ratio 20% or less?} I -- Yes --> J[Dungeon complete
pattern usage_count +1] I -- No --> K[Consider splitting off as a separate pattern] K --> B J --> L[Pattern impact tracking
one pattern edit → auto-tally of dungeons using it] L --> F classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; class L code; class F data; class J pass;

The left half of this flow (observe → decompose → map → verify → register) is the process of making a pattern; the right half (invoke → place and vary → complete → track) is the process of consuming one. Making happens rarely; consuming happens often. When the library is run well, this asymmetry turns into production efficiency.


7.3.3 The Five Base Categories

My Project A is in the action-RPG lineage, so it sorts patterns into five categories. This taxonomy is genre-dependent. A horror game would weight ambushes and narrative beats differently; a puzzle game would put environment-driven combat at the center. Don't treat the taxonomy itself as absolute — decide first what your game's core experiences are, then draw the categories.

Category Core experience Examples
pursuit Pursuit and escape Corridor pursuit, canyon escape
ambush Ambush and surprise Ambush on room entry, blind-spot ambush
puzzle_combat Environment-driven combat Levers and traps + combat
boss_phase Boss phases Boss phase 1–3 patterns
narrative_beat Narrative beats Flashback trigger, ally arrival

Within the five categories, the pattern count stays roughly between thirty and fifty. There is a reason for that number. Past a hundred patterns, a designer can no longer hold the whole library in their head. At that moment the library becomes a warehouse that takes time to search, and the designer chooses to compose from scratch instead. Once the library starts being shunned, the original goal — consistency — collapses. That is why consciously managing the cap on pattern count matters as much as category design.


7.3.4 The Format for Pinning a Pattern Down

A single pattern is pinned down as one YAML file. Below is the format actually used on Project A, anonymized. The company-specific asset names and dungeon numbers are masked, but the field structure and the way it is operated are unchanged.

---
pattern_id: pattern_corridor_pursuit_v2
category: pursuit
description: A fast enemy pursues from behind in a narrow corridor; the player makes a dodge decision at the junction
tags: [horizontal_corridor, scholar_theme_compatible]
rooms:
  - room_template: corridor_long
    size: medium
    connections_required: 2
  - room_template: junction_3way
    size: small
    connections_required: 3
npc_behaviors:
  - subtree_ref: subtree_aggressive_chase
    count: 2
  - subtree_ref: subtree_ranged_support
    count: 1
events:
  - type: trap_activation
    trigger: room_1_midpoint
  - type: enemy_spawn
    trigger: room_1_entry
difficulty_modifier: 1.2   # 1.2x load relative to a standard room
reward_modifier: 1.3
clear_time_estimate_sec: 60
art_pack_compatible: [scholar_library, generic_dungeon]
narrative_slots:
  - slot: dialogue_during_chase
    constraints: [short_dialogue, fear_emotion]
usage_count: 12            # used in 12 dungeons
last_modified: 2026-05-18
deprecated: false
---

This one file defines a piece of each of twelve dungeons at once. That is where the weight of the single line usage_count: 12 comes from. Editing this pattern means twelve dungeons are affected simultaneously — so touching a pattern file carries a different weight than fixing one room.

References like subtree_aggressive_chase and subtree_ranged_support point directly to subtrees defined in the Behavior Tree editor of 7.2. The key is that a pattern only references the BT rather than embedding it. Fix a BT, and every pattern referencing that BT follows automatically. Space (room templates) and behavior (subtrees) are managed in their own libraries; the pattern serves only as the combination table that ties the two together. Numbers like clear_time_estimate_sec and difficulty_modifier are operational values from my environment, not universal constants. Measure them yourself, with your own game's simulation and user testing, and fill them in.


7.3.5 Instantiating a Pattern into a Dungeon

When designing a dungeon, you do not compose patterns from scratch. You invoke one from the library, specify where it goes, and cover only what differs in this dungeon with overrides.

---
dungeon_id: dungeon_021_silvermark_library
pattern_instances:
  - instance: corridor_pursuit_1
    pattern_id: pattern_corridor_pursuit_v2
    placement:
      - room_id: dungeon_021_room_03
        as: corridor_long
      - room_id: dungeon_021_room_04
        as: junction_3way
    overrides:
      - field: npc_behaviors.0.subtree_ref
        value: subtree_scholar_chase   # scholar-theme variant
      - field: events.0.trigger
        value: room_1_2nd_third         # fine-tune trigger position
---

Here dungeon 021 uses the "corridor pursuit" pattern as is, but swaps the pursuing enemy from the generic one to a scholar-theme variant, and moves the point where the trap goes off from the middle of the corridor slightly back. 80% of the pattern stays; only 20% is varied.

That ratio has grounds drawn from operating experience. Too little variation (near 0%) and the dungeons feel stale, as if copied from one another. Too much (past 50%) and it is no longer the same pattern. You believe you invoked the same pattern, but the actual experience is completely different — right back to the situation of the dungeon that junior designer brought in. So we keep an operating rule: when one instance's overrides exceed half of the pattern's fields, that is not variation; it is the signal of a new pattern. It's time to split it off as its own.


7.3.6 What Shakes When You Change One Line

Edit pattern_corridor_pursuit_v2 and twelve dungeons are affected. Track that by hand and you will, without fail, miss one or two. So we keep a small tool that automatically sweeps the relationships between patterns and dungeons.

# pattern_impact.py
import json
from glob import glob

def find_dungeons_using(pattern_id):
    affected = []
    for d in glob("dungeons/*.json"):
        dungeon = json.load(open(d, encoding="utf-8"))
        for inst in dungeon.get("pattern_instances", []):
            if inst["pattern_id"] == pattern_id:
                affected.append({
                    "dungeon": dungeon["dungeon_id"],
                    "instance": inst["instance"],
                    "has_overrides": bool(inst.get("overrides")),
                })
    return affected

In the list this function returns, the key is the has_overrides flag. Dungeons without overrides use the pattern as is, so they are safe to update automatically. Dungeons with overrides may have variations of their own that collide with the pattern edit, so they need an additional human review.

Instead of a human feeling out the weight of each edit one by one, the tool reports within 5 minutes: "this edit affects 12 dungeons, and 4 of them carry variations, so look at those directly." Reducing the fear of changing a pattern is this tool's real value. When the blast radius is invisible, designers avoid touching patterns at all, and the library goes stagnant.


7.3.7 How Patterns Are Born, and Where AI Fits

Let me face the question I get most often here head-on. "Can't we just have AI write the patterns too?"

The answer is clear. No. Writing a single pattern has the designer's insight as its spine. What makes a good pursuit experience, why the junction has to be there, why the trap has to fire at the 2/3 point of the corridor rather than the middle for the tension to hold — these are the judgments of someone who has handled the game directly and watched how players respond. Have AI compose patterns from scratch, and every pattern converges to a bland, average shape. The library fills up with "patterns that aren't wrong" while "patterns that get remembered" disappear.

That doesn't mean AI has nothing to do. Of the five steps of pattern authoring, AI is a powerful assistant at two of them.

Step Output AI's role
1. Observe a good experience moment in the game Notes Designer alone
2. Decompose that moment into space, NPCs, events Draft yaml Designer alone
3. Map to existing room templates and subtrees Candidate mappings AI assist (candidate recommendation)
4. Simulation + user testing Verification AI assist (running the simulation)
5. Register in the library usage_count = 0 Designer alone

Step 3 is the heart of AI assistance. When a designer writes down a draft — "a fast enemy pursues from behind in a narrow corridor" — it is inefficient for a human to comb through room template and subtree libraries that have already grown to dozens of entries each to find what fits that intent. This is where you have AI recommend mapping candidates. Below is the prompt actually in use, anonymized.

[Input]
- The designer's draft pattern yaml (attached below)
- Room template library (list of names + tags + sizes)
- subtree library (list of names + behavior summaries)

Pick and recommend room templates and subtrees from the library to implement this pattern.
2~3 each, with a fitness score of 1~5 and a one-line rationale based on tags and size.
Don't invent anything that isn't in the library — recommendations only; if everything is 3 or below, tell me new authoring is needed.

Here is the output Claude actually returned for this prompt. I reproduce it without touching a single character, so you can see exactly where it hesitates and where it asserts.

Room template candidates:
1. corridor_long  (fitness 5)
   Rationale: tag horizontal_corridor matches; size medium fits the "narrow corridor" intent.
2. corridor_narrow_short  (fitness 3)
   Rationale: narrow, but too short to secure pursuit distance. Not enough room to place a junction.
3. hall_wide  (fitness 1)
   Rationale: a wide hall is the exact opposite of the "narrow corridor" intent. The pursuit pressure does not hold.

subtree candidates:
1. subtree_aggressive_chase  (fitness 5)
   Rationale: rear-pursuit behavior is defined, with a movement-speed weight.
2. subtree_ranged_support  (fitness 3)
   Rationale: not pursuit itself, but one unit can be placed as pressure support.
3. subtree_patrol_basic  (fitness 2)
   Rationale: patrol behavior does not match the pursuit intent. No pursuit-transition node after the trigger.

Junction room template:
I could not find an entry in the library that satisfies both the 3way junction
and the small size condition at the same time. junction_3way is the closest, but its size tag is empty,
so the fitness cannot be confirmed. Designer confirmation is needed.

The last paragraph is the most valuable part of this output. The AI said it was "not certain" about the junction room template, and it did not invent an answer that wasn't there. Receiving this report, the designer discovered that the size tag of junction_3way really was empty, and filled in that metadata first. Had the AI ignored the empty tag and assigned a fitness score of 5, the designer would have adopted it on misplaced trust.

This is the correct shape of AI assistance. The AI spreads out candidates and marks uncertainty; the choice and the responsibility stay with the designer. When every mapping comes back with low fitness, a separate task — writing a new template — is created, and that writing is once again a human's job.

[Directional signpost — compressing patterns into an "experience vector" (still premature)] Read this as a research trend, not a prescription. §7.3.1 already calls a pattern a "recipe." A pattern is close to a coordinate value — room meta, behavior subtree, event, difficulty/reward_modifier, and clear_time bundled as one. Compress that bundle into an "experience vector," and instead of combing through the flow above entry by entry when every fitness score comes back low, you could locate the new pattern as an empty region of the compressed space; the deprecated verdicts of §7.3.8 could likewise be reinforced by reading near-duplicates as coordinate distance. Three caveats attach. difficulty/reward_modifier are, as §7.3.4 says, my operational values — the axis scales differ per game, so the compressed space cannot be transplanted as is; interpolation goes only as far as "flagging" the blank, not "generating" the pattern; and actually composing the pattern on top of that flag still does not cross this section's principle that the designer's insight is the spine. The idea sits in the same spot as the dimensional-vector compression of §8.2.7, and the conceptual intuition is in Appendix M — left as territory for a team with foundations well laid to look into a few years from now.


7.3.8 Retiring Unused Patterns

A library is harder to empty than to fill. Operate one for about a year and patterns pile up that were made but almost never used. Leave them and the cost of searching the library climbs, and designers have to wade past dead options every time they pick a pattern. So we retire them regularly.

Condition Action
No usage_count increase for 6 months Classified as a deprecated candidate
Retirement decided at the review meeting Marked deprecated: true
Dungeons already using it Preserved as is (historical preservation)
New dungeons Use of that pattern prohibited

The key point is that deprecation is not deletion. Dungeons already using the pattern stay as they are. Touching dungeons running in a live service is riskier than blocking a new pattern. deprecated: true is only a sign that says "don't start using this from now on," not an order to erase the past.

Just as you pull the unused tools out of a desk drawer once a quarter and sort them, put a once-a-quarter retiring pass for the library on the calendar. Without that schedule, the library swells in one direction only, and at some point it becomes a warehouse designers turn away from.


7.3.9 Measuring the Effect Honestly

These are the changes I observed over one year of operating the pattern library on my Project A. The time figures in the table below are estimates from my environment (unverified); only the direction and the relative ratios were actually observed.

Item Before After Notes
Design time per dungeon About 2 weeks About 1 week Author's estimate; the direction is clear
Experience consistency across dungeons High variance Stable Based on user evaluation; qualitative
Average dungeons using each pattern About 8 The key metric of production efficiency
New designer onboarding About 2 months About 3 weeks Author's estimate; the biggest felt effect
Grasping the impact of a pattern change 1–2 days by hand Automated 5-minute report Effect of introducing pattern_impact.py

The most striking change is the second-to-last row: onboarding new designers. The pattern library unintentionally took on the role of a design textbook. Once a new designer could read "this is how this game builds its pursuit experience" from a single pattern file and understand it, the time a senior spent sitting beside them explaining dropped sharply. The problem of the mismatched dungeons that junior designer first brought in was, in the end, resolved by the library itself.

The figure "about 8 dungeons using each pattern on average" means the same pattern was reused eight times, and that is the honest yardstick of production efficiency. But this value of 8 is bound to my game's dungeon scale and pattern design. In a game with few dungeons, or one that demands a wholly different concept every time, this value comes out far smaller.


7.3.10 The Decision Not to Build a Library

Finally, honesty demands a point that overturns this entire chapter. A pattern library is not a cure-all. There are clearly environments where the cost of building and operating one is never recovered.

Condition Recommendation
Fewer than 5 dungeons Hands are enough; no library needed
One designer Your head is the library
Single launch, no live ops Few reuse opportunities to begin with
A completely different concept every time Low reuse ratio; ROI never recovered

The library's ROI (return on investment) is recovered only when three conditions hold together: there is live ops, there are three or more designers, and the dungeon count passes twenty. That is why a live-service MMORPG is the archetypal fit. If your project lands on any row of the table above, stop before building a library and think again. A tool has value only where there is a problem — and for a five-dungeon project, a pattern library costs more than the problem.


7.3.11 Common Failures and Remedies

Symptom Remedy
Patterns exceed 100 and designers can't memorize them Trim to 30–50; retire deprecated ones quarterly
Pattern impact tracked by hand (omissions occur) An automatic tracking tool like pattern_impact.py
Overrides exceed 80% (no real reuse) Variation too large → split into a separate pattern
Pattern authoring delegated wholesale to AI Authoring is designer insight; AI assists only at steps 3 and 4
usage_count not measured Automatic tally + review in the quarterly retrospective
No library explanation for new designers Include a library tour in onboarding materials

The second and fourth rows of this table trip teams up most often. Skip automating impact tracking and designers grow afraid of editing patterns, so the library hardens; delegate authoring to AI and the library converges to the average. Both failures kill the library's lifeblood — the reuse of verified experiences.


7.3.12 Closing Part 7

Part 7 stacked the level discipline up in three tiers. 7.1 set the standards for room metadata, tags, and connectivity (space); 7.2 covered the json-based Behavior Tree editor, subtrees, and simulation (behavior); and this chapter arrived at the pattern library, which bundles those two together with events for reuse (the operational unit). The through-line of all of Part 7 is this: in an operation that handled space and behavior separately, the same decision in the same place wobbled into a different shape every week — and that problem was solved by pinning it down in the bundle called a pattern.

This flow meshes directly with Layer integration. The vision — the spatial tone of the whole game — sits on top; below it sit the systems, the level generation rules and BT rules; rooms, BTs, and the pattern library form the content layer; dungeon instances and pattern usage statistics accumulate as data; and lint, simulation, and user telemetry verify it all at build/QA. Among these five layers, the pattern library is the spine of the content layer — and at the same time the connecting link that follows the system rules above it and generates the data statistics below it.


Key Takeaways

Next Chapter Preview


Try It Yourself — Pin Down One Pursuit Pattern

setup

  1. Create two directories, patterns/ and dungeons/, in your working folder.
  2. Prepare a file each for your room template names and subtree names, one name per line of text. (If you don't have a library yet, starting with 5 made-up names each is fine.)
  3. Save pattern_impact.py from the body text as is.

prompt

The designer writes the draft pattern yaml directly (this part is the human's job). Then hand only the mapping to the AI. Use the mapping prompt from the body text as is, attaching your own draft and the two library lists as input. Do not leave out the two key constraint lines.

- Do not invent new templates that are not in the library. Recommendations only.
- If all fitness scores are 3 or below, state explicitly that new authoring is needed.

verify

  1. Check line by line that the AI did not invent template names that are not in the library.
  2. See whether each fitness score carries a rationale based on tags and size. Do not trust a score without a rationale.
  3. Instantiate the pattern into 2 or more dungeons, run find_dungeons_using("pattern_..."), and confirm that exactly those two dungeons are caught.

Solo Scale-Down

If you are building a small game alone, a library system is overkill. Instead, pick the one dungeon segment you like most and write that experience down as a single yaml file — that alone is enough. When you build the next dungeon, open that one file, copy it, and change only 20%. The essence of a pattern library — the reuse of verified experiences — works even in a single file. When the scale grows, add the categories and the tracking tool then.

Part 8 · Balance Design

8.1 The Combat Balance Formula — The Seat of the Rulebook Called Determinism

Learning Goals for This Chapter (difficulty 🟡 practitioner · prerequisites: basic arithmetic and spreadsheet math): Separate combat balance into a formula seat and a numbers seat, and use two properties — determinism and traceability — to tell how far AI can be trusted and where humans must lock things down as a rulebook.

At 2 a.m., an alert came in: the tank class's survival rate on the live server had hit 89%. No tank was failing to finish a boss, and far too many tanks simply weren't dying. I open the data sheet looking for traces of whoever touched it. One line of the defense factor catches my eye: DEF / (DEF + 1000). Nowhere in the sheet does it say when, by whose hand, or on what grounds that 1000 came down from 1200. A hunt begins — through chat logs, through build history — that ends only when it reaches the memory of a balance designer who left the company three years ago.

Anyone who has run combat balance on a live game goes through this scene at least once. And the real cause of the scene is not that the number 1000 was wrong. It is that the number lived in the formula's seat, while the history of the formula changing existed nowhere. The combat balance formula is the domain of a game that must be the most deterministic, and the most traceable. Why these two properties become the reason AI must not be let into this seat is the spine of this chapter.

One line for non-specialists. If the z-scores, simulations, and curves in this part feel unfamiliar, that's okay. The single thing to take away is this — "Where a rule (a formula) must return the same output for the same input, AI does not go in." The judgment that separates the seats that need determinism from the seats that need exploration carries straight over to any job that handles rules that must not be wrong — accounting policies, settlement logic, contract clauses. You can take the math itself slowly, starting at 8.1.2.


8.1.1 The Formula Is a Rulebook

Do game design long enough and two kinds of documents settle into your hands: the ones that change often, and the ones that almost never do. In combat balance, the side that almost never changes is the formula. "How is damage calculated" gets touched once or twice a quarter; "what is this character's attack power" gets touched five or six times a week. Bind two flows of such different frequency into one file, and the hands that open it daily end up tearing the paper that should only open once in a while.

On Project A, which I run, combat balance is split into two seats: the formula's seat (called CombatFormula here) and the numbers' seat (CombatBalance). Here is one line that lives in the formula's seat, quoted as is.

final_damage = base_damage × dmg_multiplier × (1 − defense_factor) × variation

  base_damage    = skill_base × ATK × skill_coeff
  defense_factor = DEF / (DEF + 1000)
  variation      = uniform(0.95, 1.05)

This formula is a rulebook. Think of a board game's rules. The rules say "move as many spaces as the die shows"; they do not say "this round, if your luck is good, you may go a bit farther." Same input, same output — always. That is determinism. Put in an attack of 180, a defense of 80, and a skill coefficient of 2.1, and the same damage must come out no matter when, where, or how many times you compute it. If the same input ever produces a different output, you don't have a balance tool — you have a slot machine.

This single property, determinism, is the first reason AI must not enter this seat. We will come back to it shortly. First, let's look at what a formula has to look like to deserve the name rulebook.

The core of a combat formula does not end with one damage line. At least three lines live as a set.

# Damage
final_damage = base_damage × dmg_multiplier × (1 − defense_factor) × variation

# Critical hit
crit_damage  = final_damage × crit_multiplier
crit_chance  = base_crit + (LUK × 0.1)            # cap 50%

# Healing
heal         = base_heal × healing_power × (1 − sickness_factor)

There is a reason these three lines are written as a code block rather than prose. Natural language leaves room for interpretation. "Higher defense reduces damage" does not say whether the reduction is linear or curved, or where it stops. DEF / (DEF + 1000) reads exactly one way. A rulebook's job is to drive the room for interpretation down to zero.


8.1.2 The Curve Defines the Determinism

The single line of the defense factor, DEF / (DEF + 1000), holds this game's entire balance philosophy. Graph it and you can see why. The horizontal axis is defense; the vertical axis is the fraction of incoming damage it removes.

0% 50% ~91% Defense DEF → 0 1000 2500 5000 10000 50% damage reduction at DEF=1000 (if linear — not adopted) Steep early on Flattens late (diminishing returns)

The curve approaches its asymptote slowly. At 1000 defense it cuts damage exactly in half, and beyond that, no amount of defense ever reaches 100%. The impossibility of invincibility is built into this one line. Had it been linear, like the gray dashed line, 1000 defense would block all damage, and anything above that crosses into the nonsense territory of negative damage — healing from getting hit. That is why linear was not adopted.

Now go back to the 2 a.m. incident. Suppose someone raises this 1000 to 1200. The whole curve shifts right. The same defense now blocks less damage, so every tank in the game gets weaker and damage dealers' damage per second goes up. One constant in the formula shakes the entire game. The blast radius is nothing like changing a single number — one character's attack power. That difference is why formulas and numbers must live in separate seats, and why a formula change must always carry its history.


8.1.3 Every Formula Change Carries a History

The 2 a.m. hunt was hell for exactly one reason: there was no change history. On Project A, changing a formula is not the act of editing a line of code — it is the act of recording one decision. A separate document named CombatFormula_Decisions travels alongside the formula, and an entry reads like this.

## Decision D17 (2026-04-22)
- Change: defense_factor from DEF/(DEF+1000) → DEF/(DEF+1500)
- Reason: Tank survival rate 89% in the high-level range (LV40+) (live measurement). Cause of boss fights dragging on.
- Attempt 1: Simulated with 800 → tank death rate spiked, many full wipes within 1 minute of boss entry → rolled back
- Attempt 2: Simulated with 1200 → survival rate 75% → decent but above target (60~70%)
- Attempt 3: Adopted 1500 → simulated survival rate 65% (within target range)
- Affected atoms: combat_defense_formula, combat_tank_class_balance
- Follow-up measurement (1 week): live survival rate 67% (+2% vs. simulated prediction of 65%, within range)

This one entry answers the "why is it like this" question six months later. More important, attempts 1 and 2 survive. When the record shows why 800 failed and why 1200 was not adopted, the next person does not repeat the same mistakes. When a new balance designer joins, a stack of these decision logs is the best onboarding material there is.

One thing needs saying honestly here. The simulation figures in attempts 1–3 above (the death rates, the 75% and 65% survival rates) are the author's estimates (unverified), shown to illustrate the flow of operations. Every real game has its own curve and its own target range. But the structure — every change carries attempts, every attempt carries simulation evidence, every adoption is followed by a live measurement — is exactly how real operations run. Leave even one cell of that structure empty, and the empty cell comes back as a 2 a.m. hunt.

Here are the three seats — formula, numbers, history — at a glance.

CombatFormula Formula (rulebook) Changed 1~2 times/quarter Determinism · AI off-limits Impact: entire game CombatBalance Numbers (sheet) Changed 5~10 times/week Passes simulation gate Impact: that character _Decisions Decision history (log) 1 entry per change Reason · attempts · follow-up Key onboarding material One formula change → one decision log entry (reason · attempts · follow-up attached)

8.1.4 How One Formula Change Actually Flows

Now let's follow how D17 was decided, from the beginning. This is how a deterministic rulebook moves in practice.

flowchart TD
    A["Live measurement
Tank survival rate 89% detected"] --> B["Cause hypothesis
Defense constant 1000 overprotects late game"] B --> C["Define change candidates
1500 / 1200 / 800"] C --> D["Damage Simulator
1,000 deterministic runs per candidate"] D --> E["Result report
Survival rate, avg combat time, win rate"] E --> F{"Balance designer's call
Within target 60–70%?"} F -->|"800: death rate spikes"| G["Rejected → log attempt 1"] F -->|"1200: 75%, a bit high"| H["Held → log attempt 2"] F -->|"1500: 65%, in range"| I["Adopted → Decision D17"] I --> J["Build push (irreversible)"] J --> K["Live follow-up after 1 week
67%, +2% vs. prediction"] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; classDef fail fill:#fee2e2,stroke:#dc2626,color:#7f1d1d; class D code; class B,C,F human; class A,E,K data; class I pass; class G fail;

Look carefully at the simulator's role in this flow. The Damage Simulator runs each of the three candidates 1,000 times. Those 1,000 runs are not the same input repeated 1,000 times. The ±5% randomness of variation = uniform(0.95, 1.05) inside the formula, plus the separate randomness of critical hit chance, makes every round come out differently. You run 1,000 rounds to see the distribution: average survival rate, worst case, the spread of combat duration.

What matters is that the simulator itself must be deterministic. Given the same random seed, all 1,000 rounds must reproduce without a single digit out of place. Only then can D17's one line — "1500 produced 65%" — be reproduced and verified identically six months later. A simulator that returns different results every time turns the decision log into a lie.

I first built this damage simulator in 2008. Back then it was an Excel macro; on Project A today it is encapsulated as the balance-sim skill. In 18 years the tool's shell has changed, but the rulebook inside has never once been probabilistic. That is the point.


8.1.5 Why AI Is Strictly Off-Limits for Reward Curves and Formulas

Now we arrive at what this chapter most wants to say. At a time when AI is entering nearly every seat in game design, there is exactly one seat it must never enter: the deterministic core of combat formulas and reward curves.

An LLM is probabilistic by nature. Ask it the same question and the answer shifts a little every time. That is the source of its power for good writing and ideas, and it is fatal in the rulebook's seat. Make an LLM answer "how much damage does a character with 80 defense take?" and it may say 92 today and 94 tomorrow. That is a board game whose dice change meaning every time you turn a page of the rules.

Reward curves are more dangerous still. "XP required from level 30 to 31," once set, governs the progression speed of hundreds of thousands of players at once. Let even a ±2% wobble in, and some players grow slower than the person next to them doing the exact same hunts. Fairness collapses. Determinism is another word for fairness. So reward curves are set by human hands, entered into the sheet, and never again left to probability.

This does not mean banishing AI from the balance domain entirely. The boundary is the point.

Area AI Why
Computing damage and healing formulas Strictly off-limits Deterministic core. Break "same input = same output" and you have a slot machine
Reward and XP curves Strictly off-limits Governs hundreds of thousands of players' progression at once; any wobble collapses fairness
Simulator internals Strictly off-limits If runs can't be reproduced, the decision log becomes a lie
Detecting anomalies in simulation results Allowed z-score detection of "this character is out of range" across 1,000 results
Exploring change candidates Allowed Bounded exploration like "propose 5 candidates within base_atk ±10%"
Drafting decision log entries Allowed Meeting notes → draft Decisions entry (human review)
Summarizing follow-up reports Allowed Natural-language summaries of live data

The line is clear. AI lives only outside the deterministic core. The inside — computing and simulating — is the rulebook; the outside — analyzing, proposing, putting things into words — is AI's seat. Cross the line once, and the same input starts producing different results; from that moment, the balance tool loses trust.

This boundary has the same structure as the economy systems we will see in 8.2. There too, the resource production and consumption formulas are deterministic, and detecting inflation patterns is AI's seat. The entire balance discipline moves on the same skeleton.


8.1.6 One Step Further — A Progressive Setup Where z-Scores Propose Candidates

So far this has been the conservative application: humans create the candidates and the simulation verifies them. Go one step further, and even creating candidates can be handed to tools. The rulebook, though, still belongs to humans and to determinism.

Anomaly detection is the starting point. From the 1,000-round simulation results, look at the distribution of each character's win rate and survival rate, and measure how many standard deviations each sits from the mean — the z-score. Characters with z above 2 are flagged automatically as "outside normal range." The 2 a.m. tank would have been caught by this detection too.

For detection to lead to candidate proposals, two more things are needed. First, a defined change space. Put a column like tunable_range in the CombatBalance sheet to state explicitly within which range each number may be touched. Second, parallelized simulation. Running 10 candidates × 1,000 rounds = 10,000 rounds within the build gate window — the automated quality-gate step a build must pass before it ships — takes parallel infrastructure.

With these three in place — z-score detection, a defined change space, parallelized simulation — the decision left in the balance designer's hands narrows to one: which candidate to adopt. Creating candidates from zero and choosing among five are very different burdens. Here too, AI touches only candidate proposals and report interpretation; the computation inside the simulation and the adoption decision remain the seats of determinism and of humans.

One last point: reversibility. Editing the sheet and running the simulation are both reversible — you can undo them as much as you like. The single irreversible seat is the build push. The moment players see a number that went live, it lives on as community reaction; roll it back and the trace still remains. That is why every sign-off finishes just before the build push, while everything is still reversible.


Try It Yourself — Handling One Formula Change Safely

setup. Create a CombatFormula document that separates the combat formula from prose and records it only as code blocks, plus an empty CombatFormula_Decisions log document next to it. Split the numbers out into a separate sheet (CombatBalance).

prompt. Use AI only for analysis and drafts, never for the formula change itself. For example, hand it the simulation result CSV and ask like this.

From the attached 1,000-run simulation results, compute the z-score of each
character's win rate and organize characters with z>2 into a table. For each
character, estimate — with evidence — which number (attack/defense/skill coefficient)
is most likely the cause of the anomaly. Do not fix the numbers themselves — propose candidates only.

verify. Do not take AI's candidates on faith. Enter the candidate numbers into the CombatBalance sheet yourself, and rerun 1,000 rounds with Damage Simulator (or balance-sim) using the same seed. Check two things: (1) does the simulation result land in the target range, and (2) does one more run with the same seed reproduce it without a single digit out of place? If both pass, adopt — and the moment you adopt, write the rationale, the attempts (including rejected candidates), and the predicted values into _Decisions. One week after the build push, append the live measurement to that log.

Solo Scale-Down

Even as a solo developer with no team and no simulator, the same skeleton works. Record the formula as code blocks in code comments or a single .md file, and put a ## 변경 이력 (change history) section at the bottom of that file. Whenever you change even one formula constant, write one line with the date, the reason, and the value before the change. A 30-line Python loop is enough of a simulator. Fix the random seed, feed character numbers into the formula, run it 1,000 times, and print just the average win rate — that alone moves you from "I changed it on gut feel" to "I changed it on evidence." Use AI only to read that output CSV and summarize which characters look abnormal. The one thing you must never do, at any scale, is have an LLM compute a single line of the formula.


Key Takeaways

8.2 The Economy Model in Machinations — Catching Inflation with Simulation, Not Meetings

Primary audience: MMORPG balance/systems designers responsible for a live economy (mid-size teams of 10–50) Scaled-down version for solo/hobbyist readers: §8.2.10, "If You're Solo, Just This Much"

The first place I learned that gold was leaking was not an invoice — it was the auction house. Two months after launch, the price of enhancement stones crept up; a month later it had doubled. I called a meeting to find the cause, and everything that came out of that room was a feeling. Someone said the new dungeon rewards were too generous. Someone said hunting-ground efficiency had gone up. Someone said it was simply that we had more high-level players now. All of it was plausible, so nothing got decided. We burned an hour on guesses and ended with "let's look at more data next week."

The problem is that there is never just one resource. Gold, enhancement stones, reputation, honor, and soul stones each have sources (ways in) and sinks (ways out), and those paths feed each other. The boss that drops enhancement stones drops gold too. The gear you buy with gold burns enhancement stones. With five resources tangled into dozens of flows, no mental calculator can honestly produce even a single resource's one-week balance. This chapter is about moving that tangle into a Machinations node model, and about passing economy change decisions through a simulation gate — the economy's version of a quality gate, where the pass/fail verdict comes from a simulation run instead of meeting-room guesswork. The general theory of economy design is well covered in other books; this chapter stays focused on running that theory as an AI workflow.

Author's Note on Actual Operations The case in this chapter is an anonymized version of an economy pilot document (Economy_Machinations_Pilot) and an economy research workspace that I run in my company's R&D folder. The resource types, the source/sink structure, and the four-stage Pilot faithfully reflect actual operations; company-specific names and raw figures have been replaced for the book or stated only as ratios and directions. The AI output text is a reconstruction of real sessions.


8.2.1 An Economy Is Not 'Five Resources' but 'Dozens of Flows'

Write the economy's resources in a table and you get five rows — it looks simple. The trap is not the resources but the number of flows connecting them.

Resource source (in) sink (out)
Gold Hunting, quest rewards, auction house sales Gear purchases, enhancement, repairs, taxes
Enhancement stones Dungeon bosses, events Gear enhancement, fusion
Reputation Side quests Faction shop, class change
Honor PvP, guild wars PvP shop, guild facilities
Soul stones Boss kills Character resurrection, skill learning

Five resources — but the sources and sinks together number twenty-some, and on top of that the resources convert into one another (the auction house where gold buys enhancement stones is both a gold sink and an enhancement-stone source). The moment flows feed each other, a question like "what happens to enhancement stone prices if we release 5% more gold" has no answer from looking at one resource alone. This is where economy balance differs decisively from character balance (8.1). Character balance closes in a single formula; an economy is a dynamic system that accumulates over time — even if the one-week balance is near zero, let it accumulate for 26 weeks and the auction house collapses.

So the essence of economy work is not "picking the numbers well" but "watching, in simulation, how the flows accumulate over time." And building and revising that simulation model by hand is tedious, and something gets missed every time. Repetitive, omission-prone drafting whose review must stay firmly in human hands — this is exactly the grain of work where the division of labor between AI and humans draws itself most cleanly.

First, the skeleton of the economic loop this chapter deals with, in a single diagram.

%%{init: {"flowchart": {"defaultRenderer": "elk"}}}%%
flowchart LR
    subgraph SRC["source (resource production)"]
        H["Hunting grounds"]
        Q["Quests"]
        B["Dungeon bosses"]
        PVP["PvP/guild wars"]
    end
    subgraph POOL["pool (resource storage)"]
        G(("Gold"))
        S(("Enhancement stones"))
    end
    subgraph SINK["sink (resource consumption)"]
        UP["Gear enhancement"]
        RP["Repairs/taxes"]
        SH["Faction/guild shops"]
    end
    H --> G
    Q --> G
    B --> S
    PVP --> SH
    G -->|Auction house conversion| S
    G --> UP
    S --> UP
    G --> RP
    UP -.->|Enhancement stone demand ↑| S
    G -.->|"when net inflow > net outflow,
inflation accumulates"| POOL classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; class G,S data;

The dotted lines are the heart of this chapter. The enhancement sink pulls up demand for enhancement stones and pushes their price (UP -.-> S), and when gold's net inflow exceeds its net outflow, the surplus piles into the pool week after week and accumulates as inflation. Tracking these two dotted lines by hand calculation is impossible — which is why you need a model.


8.2.2 Machinations — A Tool That Turns an Economy into a Node Graph

Machinations is a tool that draws economic flows as a node graph and runs simulations on top of it. This is where the mermaid diagram of §8.2.1 becomes a model you can actually run.

Node Role In the diagram above
Pool Resource store Gold, enhancement stones
Source Resource production Hunting, quests, bosses
Drain Resource consumption Enhancement, repairs, shops
Converter Resource conversion Auction house (gold→enhancement stones)
Trigger Conditional activation Events, rank-up rewards

Model the economy with these nodes and run the simulation 1,000 times, and you get a distribution rather than a single result — something like "median gold price at 26 weeks: +X%; top 10% of users: +Y%." That said, Machinations is no panacea, and adopting it is itself a cost.

Limitation Remedy
Runs separately from the game code, so the two drift out of sync Calibrate monthly/quarterly against real telemetry (§8.2.6)
Readability collapses as the node graph grows Split into per-resource subgraphs; start with a single resource (§8.2.4)
The simulation uses a simplified user model Calibrate with real behavior distributions; set an error threshold
Interpreting the results depends on domain knowledge Standardize the gate that turns simulation numbers into decisions (§8.2.5)

So Machinations is not a tool you adopt unconditionally. It pays off when three conditions overlap: five or more resources + resource conversion flows + live ops. A simple economy of two or three resources is perfectly fine in Excel, and in that case the operating burden of Machinations arrives before the benefit does.


8.2.3 [Worked Transcript] Drafting a Single-Resource Gold Model with AI

A tool description alone cannot tell you what this actually produces. Here is one full cycle of moving gold — gold alone — into a Machinations model, followed from the input prompt all the way to the human veto. The input prompts can be copied and used as-is; the outputs are reconstructions of real sessions.

Stage 1 — Input: Gold Flows as a Table a Machine Can Read

First, pull gold's sources and sinks out of the data sheets into a table. This is extraction, not writing from scratch. The yaml below lists gold's three sources and four sinks; the comments note that the auction tax is the one sink that truly recovers gold, and that user-behavior distributions (kills per hour, quest completion rate) are still empty — so the AI must flag any assumptions it makes.

# gold_flows.yaml — gold single-resource flows (excerpt from the current data sheets)
resource: gold
sources:
  - id: hunting        # hunting-ground drops
    trigger: per_kill
    note: per-level-band drop curve follows the reward_curve rule
  - id: quest_reward   # quest rewards
    trigger: per_complete
  - id: market_sell    # auction house sales
    trigger: per_trade
sinks:
  - id: gear_buy       # gear purchases
  - id: enhance        # enhancement costs
  - id: repair         # repairs
  - id: tax            # auction house tax (a sink, and the key gold recovery)
# user behavior distributions (kills per hour, quest completion rate) still empty → make the AI flag any assumptions it makes

Stage 2 — The Prompt: Ask for the Model, but Force Assumptions and Format

The attached gold_flows.yaml holds the 3 sources and 4 sinks of gold as a single resource.
Draft the node specification for moving this into a Machinations model.

Rules:
1) Classify each flow by node type (Source/Drain/Pool/Converter).
2) For each node, propose a formula for 'expected weekly flow per user';
   whenever a user-behavior assumption is needed (kills per hour,
   quest completion rate, etc.), state it on its own line marked '★assumption'. Do not hide assumptions in the body.
3) Separate total source inflow from total sink outflow and compute the one-week net balance (net).
4) Reflect that the auction house tax (tax) is the only sink that permanently
   removes gold from the economy, and distinguish that the other sinks
   (enhancement, repairs) merely move gold to NPCs — not recovery.
5) ★-flag every assumption you are 'not confident' about and hand it to a human decision.

Output format: a table (node | type | weekly flow formula | assumption | confidence). No other commentary.

The prompt (in Korean above) lays down five rules: classify each flow by node type (Source/Drain/Pool/Converter); propose a formula for each node's expected weekly flow per user, putting every user-behavior assumption on its own line marked "★assumption" rather than burying it in the text; compute the one-week net balance with total source inflow and total sink outflow separated; reflect that the auction tax is the only sink that permanently removes gold from the economy, while the other sinks (enhancement, repairs) merely move gold to NPCs; and ★-flag every assumption it is not confident about, handing those to a human decision. Output format: a table only — node | type | weekly flow formula | assumption | confidence.

What actually takes up space in this prompt is not the instruction to build the model but the rules about where and how to record assumptions. Give AI free rein in economy modeling and it spits out plausible numbers with the assumptions buried. Then, when the model is wrong, you cannot point to where it is wrong. Forcing the user-behavior assumptions out of the body text and onto their own marked lines — that is where the entire weight of this request rests. An assumption you can see is an assumption you can verify.

Stage 3 — Claude's Output (Reconstructed, Excerpt)

| Node         | Type      | Weekly flow formula                     | Assumption        | Confidence |
|--------------|-----------|----------------------------------------|------------------|-------|
| hunting      | Source    | kills/h × avg drop × weekly play hours  | ★15h/wk, 80 kills/h| Mid  |
| quest_reward | Source    | daily quests × reward × 7 × completion rate | ★completion rate 0.7 | Mid |
| market_sell  | Source    | (user-to-user transfer — not net gain)  | 0 net contribution to the economy | High |
| gear_buy     | Drain     | purchase frequency × avg price           | ★1.2/wk          | Low   |
| enhance      | Drain     | enhancement attempts × cost per attempt  | ★8/wk            | Mid   |
| repair       | Drain     | hunting hours × repair rate              |                  | High  |
| tax          | Drain     | trade volume × tax rate (★only permanent recovery sink) | ★tax rate 5% | High |
| ── Net balance | net     | Σsource − Σsink ≈ +(positive for mid-activity) | with the assumptions above | ★Low |

★Assumption warning: market_sell is user↔user gold movement, contributing 0 to the economy's total.
Adding it as a source double-counts inflation. Needs separate handling in the model.
★Not confident about: the gear_buy frequency (1.2/wk) is a guess. Without real purchase telemetry
the sink may be under-/over-estimated → human decision needed.

In brief, the table classifies the seven nodes and proposes weekly flow formulas with ★ assumptions — 15 hours of play per week, 80 kills per hour, a 0.7 quest completion rate, 1.2 gear purchases per week, 8 enhancement attempts per week, a 5% auction tax — and marks the net balance itself as low-confidence. It closes with two warnings: that market_sell is user-to-user gold movement contributing zero to the economy's total (adding it as a source double-counts inflation), and that the gear_buy frequency is a guess that needs a human decision.

The most valuable part of the output is not the table but the two lines at the bottom — the "★가정 경고" (assumption warning) and "★확신 못 하는 점" (points it is not confident about). The AI reported two weaknesses in its own model on its own. A good prompt makes the AI say, "I don't trust this assumption."

Stage 4 — Verification and Veto (The Human's Seat)

Do not feed this output into the model as-is. One of the two ★ flags the AI reported was, in fact, a defect that would have broken the model.

The AI initially classified market_sell (auction house sales) as a Source. But an auction house sale is user A's gold transferring to user B — it does not create new gold in the economy. Add it to source inflow and you count inflation twice. The AI did flag this itself in the ★ assumption warning, yet in the table body it still left the node sitting in the Source column — it reported the problem without removing it from the model: an output that was only half right. This was also partly a data defect on the human side: the input yaml never specified the nature of market_sell (user-to-user transfer vs. new creation).

So I send a follow-up request.

market_sell is a user↔user gold transfer, not an economy-level source (fixing
an input omission). Remove this node from the source total and instead reflect
it in the model only as 'the auction house tax (tax) permanently recovering a
fraction of the transferred amount'. Recompute the net balance, and show in one line how excluding market_sell changed net.

(The follow-up, in Korean above, says: market_sell is a user-to-user gold transfer, not an economy-level source — an input omission, now corrected. Remove the node from the source total and reflect it in the model only as "the auction tax permanently recovering a fraction of the transferred amount." Recompute the net balance, and show in one line how excluding market_sell changed it.)

The AI answered with a revised model: market_sell removed from the sources, only the tax kept as a sink. The net balance came out lower than the first estimate — revealing that with auction house sales wrongly counted as a source, we had been overestimating inflation. This one round trip is the whole point. Build the model by hand from scratch and it takes half a day, and the person who made the node-classification mistake is the worst-placed person to catch it; with an AI draft + forced assumption disclosure + one veto, it is under an hour, and because a human adjudicates the ★ flags the AI itself reported, defects like double-counting get caught before they ever enter the model (author's estimate — the time saved varies with team and resource count, so read this less as absolute values and more as the structural difference between "by hand from scratch" and "draft + review").


8.2.4 One Resource at a Time — Rolling It In Through a Four-Stage Pilot

Closing one gold model does not license modeling every resource at once. My own rollout did not load everything in one go either. It followed four stages: start from a single resource, pass verification and calibration, then expand.

Stage Scope Key gate
1. Single-resource (gold) modeling 3 sources, 4 sinks; the §8.2.3 session Node classification, explicit assumptions
2. Simulation vs. reality One simulated week of net vs. one week of telemetry Pass/fail against the error threshold
3. Model precision calibration Add user behavior distributions (low/mid/high activity) Re-measure error per segment
4. Resource expansion (all five) Add enhancement stones, reputation, honor, soul stones in stages Verify conversion flows (auction house)

The comparison in stage 2 is the heart of these four stages. When simulation and reality disagree, what is wrong is the model, not the game. Make decisions on a misaligned model and those decisions come back as live incidents. So expansion (stage 4) happens only after the verification of stages 2 and 3 has passed. Break this order — skip single-resource verification and load all five at once — and you can no longer even isolate which resource's model is wrong.


8.2.5 The Simulation Gate — Putting a Barrier in Front of Economy Change Decisions

Once the model passes verification, stand a simulation gate in front of every change decision that affects the economy. This is where decisions that used to pass on meeting-room "feel" start passing on simulation results instead.

Decision type Simulation required
Adding a new source or sink Mandatory
Changing resource conversion rates (auction house exchange rates, etc.) Mandatory
Designing rewards for a new dungeon or event Mandatory
Price changes (±10% or more) Mandatory
Verifying a new class's efficiency Mandatory
UI changes and other non-economy work Exempt

To see how the gate actually works, here is one decision passing through it, on the gold model verified in §8.2.3.

[Simulation Gate — Event Reward Decision] (Reconstruction of the Actual Format)

[Proposal]   Weekend event: daily login reward +500 gold
[Gate]       new source added → simulation mandatory
[Sim results, 1000 runs]
  - weekly gold net balance: +6,900 → +10,400 (+50%)
  - 26-week cumulative median gold price ~+28% (inflation warning: exceeds ±10%)
  - top 10% active users: ~+41% (large segment variance)
[Verdict]    FAIL — exceeds the stable range (±10%/long-term)
[Correction] attach a simultaneous sink to the event source: event-limited shop (gold recovery)
             re-simulate → 26-week cumulative +9% (PASS)

(The gate record, in Korean above: a weekend event proposing +500 gold in daily login rewards is a new source, so simulation is mandatory. Over 1,000 runs, the weekly gold net balance jumps from +6,900 to +10,400 (+50%); over 26 weeks the median gold price accumulates roughly +28% — an inflation warning, past the ±10% line — and the top 10% of active users reach about +41%. Verdict: FAIL. The corrected proposal attaches a simultaneous sink — an event-limited shop that recovers gold — and re-simulates at +9% over 26 weeks: PASS.)

The value of the gate is in the last two lines. If "let's hand out +500" had been a meeting-room guess, it would have passed as "probably fine." The simulation gate converts that decision into 26 weeks of +28% inflation and shows it to you, and it forces the correction: if you add a source, attach a sink along with it. Judging economy changes by simulation pass/fail instead of guesswork — that is the entire gate.

One trap worth flagging here, because people fall into it constantly: segment variance. Even when the mid-activity user sits at +28%, the top 10% sits at +41%. The users who earn the most gold accumulate inflation the fastest, so the simulation must run per segment, not just on the average. Look only at the average and you miss the price collapse driven by high-activity users.


8.2.6 The Model Evolves Monthly on Post-Launch Telemetry

For the simulation gate to be trusted, the model must not drift from the actual game. The game changes every week, so the model has to be recalibrated to keep up. After launch, check the model against real telemetry monthly (quarterly in periods with few changes).

Model calibration cycle (monthly)
─────────────────────────────────
1. Extract one month of real user telemetry (flows aggregated per resource)
2. Compute measured source/sink flows per segment (low/mid/high activity)
3. Compare item by item against the Machinations simulation
4. Items with error >15% = adjust model parameters (that item's ★assumption was wrong)
5. Re-simulate after adjustment → use as next month's gate baseline model

(The monthly cycle, in Korean above: extract one month of real user telemetry, aggregated as flows per resource; compute measured source/sink flows per segment — low/mid/high activity; compare item by item against the Machinations simulation; for any item with error >15%, adjust the model parameters — that item's ★ assumption was wrong; re-simulate and use the adjusted model as next month's gate baseline.)

Step 4 is the core. A high-error item is a signal that one of the "★ assumptions it was not confident about," which the AI reported back in §8.2.3, has diverged from reality. For instance, if the AI's guessed gear_buy frequency (1.2 per week) measures at 2 per week in telemetry, replace the assumption with the telemetry value. Stop this calibration and the model drifts slowly away from the game, until one quarter the simulation gate "passed something that inflated live anyway." At that moment, trust in the simulation itself collapses retroactively. Calibration is not a side chore of operations — it is the regular cycle that keeps the gate alive.


8.2.7 Progressive Application — Anomaly Detection, Change Spaces, Parallel Simulation

Everything up to here is the 'conservative application' of economy modeling: a human proposes a change, the model verifies it, and the result decides. One step further, the three axes of progressive application seen in 8.1.6 — z-score detection, defined change spaces, parallel simulation — open up on this economic infrastructure in exactly the same way.

First, anomaly detection. Instead of a human eyeballing the error comparison in the monthly calibration cycle (§8.2.6), code picks out and surfaces the items whose model-vs-measured deviation crosses a threshold. You no longer learn that enhancement stone prices doubled by watching the auction house; an alert saying "enhancement stone source flow has deviated +30% from the model" arrives before the meeting does.

Second, defining a change space. Instead of the binary "release +500 in rewards / don't," define the change space — a reward range (0 to +1000) and a range of accompanying sinks — and you can search that space for combinations that keep inflation within ±10%. Humans set the "from where to where"; the search for the optimal combination inside it is automated.

Third, parallel simulation. Instead of running one proposal 1,000 times, run dozens of candidates from the change space in parallel, 1,000 runs each, and compare the distributions in a single pass. The meeting room where proposals were debated one at a time becomes a comparison of simulation results across a candidate matrix.

The common idea is to move the seat where a human proposes changes to one where code searches a change space. But this is strictly for after the conservative application (§8.2.3–8.2.6) runs stably and the model has been verified against telemetry. Auto-search a change space with an unverified model, and a wrong model will confidently hand you a wrong optimum.

[Radical Application — Compressing the Economy into a 'Dimension Vector' and Searching There] (Still Premature)

This goes one step beyond progressive application. Read it as a research direction, not a claim (if dimension vectors and embeddings are new to you, look first at the one-page 'map' in Appendix M — all five 'signposts' in this book run on top of that picture, and what follows will read easily). The economy model we have built to this point is a high-complexity system — five resources tangled in dozens of flows — and even the change-space search of §8.2.7 ultimately works by pinning each of those dozens of flows as a parameter and iterating. The radical idea is to compress that complexity itself into a dimension vector, then look for solutions in the compressed space.

A seemingly distant analogy supplies the clue. Cooking recipes — a domain usually held up as qualitative and ungraspable — were handled in one study (Epicure — Radzikowski & Chen, 2026, arXiv:2605.22391; demo at epicure.kaikaku.ai) by distilling 1,790 standard ingredients from 4.14 million recipes across 11 sources and compressing the relationships between ingredients into vectors of several hundred dimensions. The point is that even a qualitative target like "taste," once the relationships between ingredients are converted into coordinates, makes similar recipes cluster close together in vector space — and you can interpolate between them to search for new combinations. Epicure itself demonstrates this kind of interpolated search in the compressed space, rotating one ingredient toward a particular cuisine to find its counterpart.

The economy works on the same principle. Express the economy's state as a vector whose dimensions are its sources, sinks, and conversion flows, and "a stable economy within ±10% inflation" becomes a region of that space. Then, instead of putting proposals through the simulator one at a time, a path opens to search for solutions directly in or near that stable region. The parallel simulation of §8.2.7, which compares candidates by running every one of them, could narrow down to a single search over the compressed space.

Why "still premature"? First, deciding what to take as a dimension (which flows are independent and which are dependent) is itself a hard domain problem. Second, compression by its nature throws information away, and a live incident can erupt in a dimension you threw away. Third, all of this means anything only when the conservative application's telemetry verification (§8.2.6) is solid — if the pre-compression model is misaligned with the game, compression will simply compress that error neatly along with everything else. So this section is a signpost, not a prescription. The work for today is to run the conservative application honestly; dimension vectors remain a research area for teams with that foundation to look into a few years from now.


8.2.8 Measurement — Where Meetings Gave Way to Simulation

Here is before-and-after, around the tool's adoption. The times and frequencies below capture the direction we felt in early operations; read them for which way things moved, not as precise absolute values.

Item Before (meetings, hand calculation) After (simulation gate)
Economy change decision → applied 2–4 weeks (guessing, repeated re-discussion) 1–3 days (one simulation verification)
Inflation incidents 1–2 per quarter (found after the fact) 0–1 per quarter (blocked in advance by the gate)
Frequency of adding sources/sinks 1–2 per quarter (conservative out of fear) 1–2 per month (simulation guarantees safety)
Economy meeting frequency 3–4 per week 1–2 per week

The last row means more than the numbers. Meeting frequency dropped because simulation replaced debate. When "I think enhancement stones are headed for inflation" becomes "the simulation says +28% over 26 weeks," the hour spent arguing over guesses becomes a five-minute results share. This is exactly the concept codified in my own system's retrospectives (the atom automation_signal_value_over_time_savings — the value of automation is signal exposure, not time saved). The real output of the simulation gate is not the hours saved; it is that the seat guesswork used to occupy in meetings is now occupied by numbers.

One thing I keep honest, though. The "1–2 per quarter → 0–1" row is not a precise measurement; it is the direction of an operational impression. What counts as an inflation incident depends on definition (how many ± percent of price movement you call an incident), so read it less as absolute counts and more as the structural shift from "found after the fact" to "blocked in advance."


8.2.9 Common Failures

Pattern Why it fails Remedy
Modeling every resource at once Cannot isolate which resource's model is wrong Start with a single-resource Pilot (§8.2.4)
Accepting the AI model's assumptions without review Defects like auction house double-counting go straight in Force explicit assumptions + human veto (§8.2.3)
Economy changes without a simulation gate Recovering from inflation after the fact costs a fortune Define mandatory simulation items (§8.2.5)
Simulating only the average, ignoring segments Misses the price collapse driven by high-activity users Per-segment simulation (§8.2.5)
No post-launch telemetry calibration Model drifts from the game; trust in the gate collapses Monthly/quarterly calibration cycle (§8.2.6)
Auto-searching a change space with an unverified model A wrong model confidently produces a wrong optimum Progressive only after conservative is stable (§8.2.7)

The second one is missed most often. As with the auction house double-counting in §8.2.3, AI will confidently produce a plausible model while merely ★-flagging the weak points of its own assumptions and leaving the defect in the body. If no human adjudicates those ★ flags, the wrong model passes — and every simulation decision built on top of it is wrong together.


8.2.10 Try It Yourself — One Step You Can Take Today

If You're Solo, Just This Much: You don't need Machinations or telemetry. Pick one resource from your own game (or a game you love), write its sources and sinks on paper, paste in the prompt from §8.2.3 as-is, and get a draft one-week net balance model. Then pick one of the assumptions the AI marked with ★ and push back: "I don't trust this assumption — justify it again." You will feel, firsthand, what an economy model really is — a bundle of assumptions — and how the conclusion flips when one of those assumptions is wrong.

If you're on a team, start with this one step. Not every resource — pick the single most troublesome resource (usually gold or enhancement stones), build only the single-resource model of §8.2.3 first, and put the simulation gate of §8.2.5 on one kind of economy change decision (say, event rewards). One resource plus one decision type is already enough to replace the argument over guesses with a single line of numbers.

Summed up as setup → prompt → verify — setup: extract one problem resource's sources and sinks into yaml. prompt: request a node model draft in the §8.2.3 format, forcing user-behavior assumptions to be marked with ★. verify: a human personally vetoes and re-requests against the ★ assumptions the AI reported and the node classifications (especially user-to-user transfer vs. new creation).


Key Takeaways

Next Chapter Preview

8.3 Damage Simulator — The Day Spec DPS and Sim Output Diverged

One early morning in 2008, I sat in front of a single Excel sheet, checking the same number for the third time. The game design document (GDD) listed a certain sword character's damage per second (DPS) as 847. But the simulator I had run for the first time that day, fed the same character with the same stats, spat out 612. A 27% difference. One of the two was lying, and I didn't yet know which.

Spec DPS is a promise on paper: arithmetic that multiplies one skill hit's damage by how often it fires. Simulator DPS is what you get when you actually swing that promise 1,000 times. Cooldowns overlap, cast animations eat up time, critical hits land less often than the expected value says — friction the paper knows nothing about creeps in. That 27% gap is exactly where a balance designer makes a living. Trust the paper and you cry after launch.

This chapter is the story of that one tool: the Damage Simulator I built in 2008 and have not put down since. How I tracked the exact point where spec and output diverge, and how, 18 years later, I attached AI to that tracking — followed through a single real worked transcript.


8.3.1 Why Spec DPS Always Lies

First, let's dissect what that 612 versus 847 really was. The junior designer who wrote the spec (Team Member A from here on) had done nothing wrong. He multiplied exactly what the skill table said.

The spec-side DPS calculation looks like this. Assume a character with three skills.

Skill Single-Hit Damage Cooldown Cast Time
Horizontal Slash (횡베기) 320 3.0s 0.6s
Thrust (찌르기) 540 6.0s 0.9s
Basic Attack (평타) 180 1.2s 0.4s

Team Member A's spec calculation rested on the ideal assumption that every skill is used the moment its cooldown comes back. Horizontal Slash deals 320 every 3 seconds, Thrust deals 540 every 6 seconds, and Basic Attack fills the empty time. The arithmetic comes out to a clean 847. On paper, the character has more than one pair of hands, so cast animations never block one another.

The simulator returned 612 for exactly one reason: the character has one pair of hands. While the 0.9-second Thrust is being cast, Horizontal Slash cannot fire even when its cooldown is up. This global cooldown (GCD) collision — cast animations devouring each other — does not exist in the spec. Paper is a frictionless vacuum; the sim is a battlefield with friction. The first truth I learned in 24 years of doing this work is this: spec DPS is an upper bound, not a measurement. And players play the game at the measurement, not the upper bound.


8.3.2 Worked Transcript — Grilling the AI About Where 612 Came From

In 2008, I closed that 27% gap by hand. I read the sim log line by line with my own eyes, counting the frames where skills got blocked. It took half a day. Eighteen years later, I throw the same log at Claude. Below is the full text of an actual session. I have not summarized it, and I include the first attempt that fell short — because that falling short is half of how to use this tool.

Here is the prompt I threw.

Attached is part of a single-character 1,000-tick combat log from our Damage Simulator (damage_sim/, Python). Per the spec, this character's DPS is 847, but the sim outputs 612. Find the cause of the 27% gap in the log. The log format is [tick] event skill=SKILL state=STATE. No guessing — cite only evidence that is in the log. If the evidence is insufficient, say it is insufficient.

[0.0] cast skill=찌르기 state=START
[0.9] cast skill=찌르기 state=HIT dmg=540
[0.9] cast skill=횡베기 state=BLOCKED reason=GCD
[1.5] cast skill=평타 state=HIT dmg=180
[1.9] cast skill=평타 state=BLOCKED reason=GCD
[3.0] cast skill=횡베기 state=HIT dmg=320
[3.6] cast skill=평타 state=HIT dmg=180
... (994 lines omitted)
[997.2] cast skill=찌르기 state=HIT dmg=540

Claude's raw output looked like this (untouched).

Facts confirmed in the log:

  1. state=BLOCKED reason=GCD events appear repeatedly. At [0.9], Horizontal Slash (횡베기) was blocked by the GCD (global cooldown); at [1.9], Basic Attack (평타) was.
  2. Until the Thrust (찌르기) cast finishes (0.0→0.9, 0.9s), no other skill can fire. The cast time appears to occupy the GCD.
  3. Therefore the spec's assumption — "every skill fires on every cooldown" — breaks down. In practice, cast animations collide and some activations are lost.

However, the provided log is only 7 lines out of 1,000. Whether the exact 27% gap is explained by GCD collisions alone cannot be concluded from these 7 lines. To verify the 27%, we need the total count of BLOCKED events and the total damage lost because of them. If you provide the BLOCKED aggregation for the full log, I will confirm it quantitatively.

The stop here is what matters. The AI pointed at the cause but did not prove the 27%, and it stated that limit itself. Where a paper-only designer would have ended the meeting with "so it's the GCD," the tool held its ground: "seven lines aren't enough to know." I judged that demand legitimate — this is what human verification is. So I pulled the BLOCKED aggregation for the full log and threw it back.

Here is the full 1,000-tick aggregation. BLOCKED total: 412, of which 388 were due to GCD and 24 due to cooldown not yet recharged. Estimated damage lost to BLOCKED is about 19,800 in Horizontal Slash equivalents. Over the full 1,000-tick (= second) fight, spec-expected damage is 847,000 and sim-measured cumulative damage is 612,300. Re-examine the 27% gap with these numbers.

Claude's second output.

Verifying with the aggregates.

Composition of the gap: - Directly lost to GCD collisions (388 events): the estimate of about 19,800 is in Horizontal Slash equivalents, but you also need to add the opportunity cost — higher-damage skills could have filled the lost slots. Converting the lost slots at average damage makes the loss larger than 19,800. - However, BLOCKED alone does not explain the full 234,700. The interpretation of the remaining gap changes depending on whether the spec's 847 "already folds the critical-hit expectation into the average" or is "a base value without crits." Tell me the spec's crit assumption behind 847 and I will decompose the residual gap.

It stopped again. And again it stopped in exactly the right place. How the spec's 847 handled critical hits — that information was not in the log; it existed only inside Team Member A's head. The AI knew the boundary of the data it had. When I checked, the spec's 847 had flattened a 20% crit chance at 1.5x into the expected value, and the sim was one run, seed fixed, in which the actual crit count came in below expectation. In other words, the 27% gap was a mix of GCD collision (structural, reproducible) and crit variance (statistical, this one run's luck).

That decomposition is the conclusion. The GCD-collision share is a real problem that must be fixed in the design; the crit-variance share is noise that disappears once you change seeds and average 1,000 runs. Mix the two together, decide "the character is weak," and hand out a buff — and a character that was fine on the 1,000-run average becomes overpowered. The paper didn't know it, a single sim run didn't know it, and the AI on its own didn't know it either — what produced this distinction was a human verification that brought the log aggregation and the spec's hidden assumption to the table.


8.3.3 One Set of Inputs, One Set of Outputs — Anatomy of the Sim

Let's lay out, as one set, the inputs and outputs of the tool that session was peering into. A simulator is, in the end, an honest function. Same input, same output. The input gathers from three sources.

flowchart LR
    A["Game data sheets
(skill table, character stats)
read-only"] --> SIM B["Scenario yaml
(combat conditions, duration,
target composition)"] --> SIM C["seed=42
(fixed RNG)"] --> SIM SIM["Damage Simulator
domain/formulas.py
run_combat() × 1000 ticks"] SIM --> R1["Measured cumulative damage
612,300"] SIM --> R2["BLOCKED log
412 events (GCD 388)"] SIM --> R3["Critical hits
measured 17.2% vs expected 20%"] R1 --> REP["markdown report
spec 847 vs sim 612
gap decomposition: structural 19% + variance 8%"] R2 --> REP R3 --> REP classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; class SIM code; class A,B,C,R1,R2,R3,REP data;

The point of this diagram is that the arrows run one way. The game data sheets are only ever read by the simulator. The sim never writes the data back. The rule that prevented the most accidents over 18 years was the direction of that single arrow. The moment a sim starts keeping its own copy of the data, then the day after the game data changes, the sim is simulating yesterday's world. Hold a meeting over a report produced that way, and the entire meeting ends up arguing about yesterday's world.

A concrete input set (the scenario yaml) looks like this.

# scenarios/single_dps_check.yaml
scenario: single_target_dps
duration_ticks: 1000      # assume 1 tick = 0.1s, 100s combat
seed: 42                  # determinism — same input, same output
actor:
  char_id: K_004          # read from game data sheets
  skill_rotation: optimal # on GCD collision, prefer highest expected damage
target:
  defense: 1200
  hp: infinite            # infinite-HP dummy for DPS measurement
report:
  compare_to_spec: 847    # feed spec DPS to auto-decompose the gap

And one output set (report excerpt) looks like this.

# Damage Simulator Report — K_004 single DPS
Input: scenarios/single_dps_check.yaml | seed=42 | data rev. 2026-06-05

## Versus Spec
- Spec DPS:          847   (20% crit · 1.5x expectation flattened in)
- Sim measured DPS:  612   (this seed, 1 run)
- Gap:              -27.7%

## Gap Decomposition
- Structural (GCD collision, reproduced):  -19.2%  ← design review target
- Statistical (crit variance, this run):   -8.5%  ← expected to vanish on 1000-run average

## Reproduction Verification
- seed=42 re-run 3 times → 612,300 / 612,300 / 612,300 (identical)
- seed 0~999, 1000-run average DPS → 731 (after crit variance removed)

Look at the last line. The average over 1,000 runs with seeds 0–999 was 731. The gap between the spec's 847 and the 1,000-run average of 731 — 116 (13.7%) — is the size of the real structural problem, the GCD collision. That 731, not the single run's 612, is what should enter the design meeting as input. Not the paper's 847, not the unlucky single run's 612, but the 731 that 1,000 runs agreed on. Getting that number into your hand is the balance designer's job.


8.3.4 The Hands of 2008 and the Hands of 2026

This tool has lived 18 years, but not as the same code. I kept the hanger and changed the clothes on it five times. The hanger is the logic of the report above — split spec from measurement, decompose the gap into structure and variance, verify by reproduction. That procedure is, word for word, the same in 2008 Excel VBA (Excel's macro language) as in 2026 Python.

Period Clothes (Tech) Hanger (Unchanged Procedure)
2008–2011 Excel VBA, 1:1 Spec-versus-measured gap decomposition
2012–2016 C# console, N:N Same
2017–2020 Python + Web Same
2021–2024 Python + ML Same (+ player distribution modeled)
2025– Python + LLM assist Same (+ log queries, hypothesis generation)

The secret to surviving five changes of clothes is carved into the folder structure.

damage_sim/
├── domain/          # the hanger — unchanged for 18 years
│   ├── formulas.py      # damage formulas · GCD collision checks
│   └── metrics.py       # gap decomposition logic
├── adapters/        # game data read-only
│   └── excel_reader.py
├── runners/         # clothes — replaced whenever the tech changes
│   └── cli_runner.py
└── reporters/       # clothes — report output format
    └── markdown_report.py

When the tech changes, only runners/ and reporters/ get rewritten. The gap-decomposition logic in domain/ survives intact as an 18-year asset. The GCD collision check I once wrote in Excel cells in 2008 is running in today's formulas.py with nothing changed but the function signature. Nail a tool to one technology and it grows old and dies with that technology — I learned that by burying several dead tools.

The LLM I attached in 2025 is not a new hanger; it is a new pair of hands. As the session above showed, AI is the hand that reads logs and forms hypotheses, and log tracing that used to take half a day now takes minutes. But it does not touch the hanger — whether the gap is 27%, what percentage of hits crit: those are still decided by the seed-fixed deterministic core. The moment an LLM steps into that seat, regression verification becomes impossible and the tool dies.


8.3.5 The Deterministic Core and What Lies Outside — Where to Draw the Line

If I had to draw only one line in a balance tool, I would draw it at the boundary of the deterministic core. Inside, same input must yield same output with steel-grade certainty; outside, humans and AI are free to throw hypotheses around.

Inside (deterministic — AI forbidden): - Damage formulas, GCD collision checks, critical-hit rolls, cumulative aggregation. - Run it three times with seed=42 and 612,300 must come out three times, identical. If that breaks, yesterday's report and today's report can no longer be compared.

Outside (hypothesis and interpretation — AI welcome): - Causal questions like "why is this character combo's win rate abnormal." - Finding BLOCKED patterns in logs, drafting natural-language reports, drafting scenario yaml.

The worked transcript above moved exactly along this line. The AI, on the outside, was quick to form the hypothesis that GCD collisions were the cause. But the number 27%, the number 612 — those were values computed by the deterministic core to the very end, and the AI only took those values and interpreted them. And twice it stopped, saying "this data is not enough to conclude" — demanding information the deterministic core could not supply (the spec's crit assumption). That stopping is the mark of a good tool: not mistaking a hypothesis for a diagnosis.

One honest disclosure about the numbers. The concrete figures in this chapter — 847, 612, 731, 412 events — are example values constructed for explanation. But the direction (spec DPS always comes out above the sim measurement), the structure (the gap decomposes into structural collision and statistical variance), and the principle (a fixed seed is the precondition for regression verification) are things I have confirmed repeatedly across 18 years of actually operating this tool since 2008. The size of the ratios varies by project; the direction and the structure have not changed.


Try It Yourself — One Round of Spec-Versus-Sim Gap Decomposition

setup. Pick one character from your game data and gather its skill table (damage, cooldown, cast time) and its spec DPS. If you have no simulator, write a minimal script that runs a 1,000-tick single-target fight. The one thing that matters: it must take seed as an argument so the run can be pinned.

prompt. Throw the sim log (including BLOCKED events) and the spec DPS together.

Attached are a single character's 1,000-tick combat log and the BLOCKED aggregation. Spec DPS is [N] but the sim outputs [M]. Decompose the cause of the gap using log evidence only. Separate structural causes (reproducible collisions) from statistical causes (this run's variance). If the evidence is insufficient, say so and name what else you need.

verify. Take the structural cause the AI pointed at and verify it on a 1,000-run average with varied seeds. If the gap survives the average, it is a real structural problem; if it vanishes, it was variance noise. If the AI stops and says "I can't conclude," that is not a failure — it is normal. A human steps into the spot where it stopped and fills in the spec's hidden assumptions.

Solo Scale-Down

If you are a solo developer with no simulator and no ML, you can run the same procedure with one Excel sheet and AI. Put the skill table on a sheet and build a 1,000-row sim in a column, rolling crits with RAND(). You can't fix the seed, so press F9 to recalculate 100 times and read the average by eye. Throw that average and the spec DPS at the AI and ask it to "split the gap into structural causes and variance causes." The tool is small, but the hanger — spec-versus-measured gap decomposition, separating structure from variance, verification by reproduction — stands just the same.


Key Takeaways

Next Chapter Preview

8.4 AI-Assisted Balance Simulation

Friday, 4 p.m. The alpha build's 5:5 PvP auto-simulation finished its 1,200 matches. The result JSON was 4 megabytes. Somewhere inside it sat a single line that read "Team A win rate 92%" — while the overall average win rate was 52%. I spent 40 minutes finding that line, and I clocked out without ever learning why.

Balance is the domain of determinism. Put the same input through the same formula and you always get the same damage. That is why a damage simulator must be code, and why a reward curve must be drawn by a human hand — this is one place AI must not set foot. But the periphery of that deterministic core — finding the one strange line in 1,200 match results, forming hypotheses about why, shortlisting what to change, and running those candidates back through the sim — that peripheral labor eats most of a balancer's day. This chapter is about attaching AI to that periphery. With the core left untouched.

8.4.1 The Core Is Code, the Periphery Is Human Labor

That 2008-vintage damage simulator from 8.3 — the one whose deterministic core survived three changes of engine and company — is this chapter's starting point. Same input, same output: that property is the entire trust a balance tool has. If you run the same build twice and get different win rates, the tool belongs in the trash.

So when you sketch the skeleton of balance work, you get a deterministic block in the middle, with human handwork hanging off its entrance and its exit. Below is that skeleton broken down — the deterministic region (blue) and the region where humans and AI step in (orange), split by color.

Deterministic simulation simulate_dps() input=output, no AI Slot 1 Scenario generation Slot 2 Change candidate search Slot 3 Reports Slot 4 Anomaly interpretation Slot 5 Action proposals Balancer (adopt/reject)

Only the blue box in the middle is code. The other five orange boxes are all human labor — judging, interpreting, writing — and those five spots are the only places AI can enter. The moment you ask a large language model (LLM) to "calculate this character's damage per second (DPS)," non-determinism — different numbers from the same input — leaks into the core, and the tool loses its credibility in less than 18 days.

The spine of this chapter is therefore simple: keep the core as code to the very end, attach AI to the five slots at the entrance and exit, and automate the exit side first — the most labor-hungry part, finding the one strange line in 1,200 match results and forming hypotheses about it.

8.4.2 Worked Transcript: Tracking Down the 92% Win Rate Line

Back to that 92% from the opening. This time, instead of a human wandering for 40 minutes, a deterministic detector picks out the line, an LLM forms the hypotheses, and the sim verifies them — and we follow one full cycle from start to finish. Nothing is summarized; the raw output the tools actually produced stays as is.

Step 1 — Anomaly Detection Is Done by Code (z-score)

Picking the "strange" matches out of 1,200 results is a job for statistics, not the LLM. Compute each metric's mean and standard deviation, then split by how many standard deviations a value sits from the mean — its z-score. Past the threshold, it's an outlier. This is deterministic; hallucination has no way in.

def find_outliers(results, threshold=2.5):
    # results: list of {metric_name: value} dicts, one per sim match
    means, stds = compute_per_metric(results)   # per-metric mean / std dev
    outliers = []
    for r in results:
        for metric, value in r.items():
            if stds[metric] == 0:               # zero variance → not comparable, skip
                continue
            z = abs(value - means[metric]) / stds[metric]
            if z > threshold:
                outliers.append((r["scenario_id"], metric, value, round(z, 2)))
    return sorted(outliers, key=lambda x: -x[3])  # largest z first

Running it yields the following — of 1,200 matches, only 3 crossed the 2.5 threshold.

[("pvp_5v5_S0417", "team_a_winrate", 0.92, 4.1),
 ("pvp_5v5_S0417", "match_duration",  41.0, 2.9),
 ("pvp_5v5_S0822", "team_b_winrate", 0.18, 2.6)]

The first line, the one with the largest z — scenario pvp_5v5_S0417, win rate 0.92 (z=4.1) — is the very line I spent 40 minutes hunting in the opening. This is not work for a human eyeballing a 4-megabyte JSON; statistics narrows it down to 3 cases. Everything up to here is the core; everything from here on is AI.

Step 2 — The LLM Forms Hypotheses (No Definitive Diagnoses)

Now the line goes to the LLM. But not as "diagnose the cause." The LLM only throws out a handful of possible cause hypotheses from domain knowledge; which one is real gets decided by the sim again. Here is the full prompt.

[outlier]
Scenario: pvp_5v5_S0417 — 5:5 PvP
Team A composition: [refgame_archer_07, refgame_archer_07, refgame_archer_07,
            refgame_hybrid_21, refgame_hybrid_21]
Metric: team_a_winrate 0.92 (overall average 0.52, z = 4.1)
Secondary metric: match_duration 41.0s (average 28s, z = 2.9)

[related info]
- refgame_archer_07: ranged support, skill "표식" (Mark) — debuff: target takes +12% more damage
- refgame_hybrid_21: melee/ranged hybrid dealer, skill "추격" (Pursuit) — +20% move speed while chasing a debuffed target
- These two characters have almost never been placed on the same team (combination frequency: bottom 5%)

Request:
Present 3–5 possible cause hypotheses for this outlier.
- One extra verification sim line for each hypothesis
- No definitive diagnoses. Hypotheses only, at the level of "it may be ...".

Response format: numbered list. Hypotheses only — no preamble, no closing.

Here is Claude's raw output, carried over verbatim (not summarized, not tidied).

1. The "표식" (Mark) debuff and the "추격" (Pursuit) move-speed buff may have
   stacked multiplicatively: 3x archer_07 blanket Marks → hybrid_21 chases
   without interruption, forming a positive feedback loop.
   Verify: replace hybrid_21 with a plain melee dealer and re-sim the same scenario.

2. The Mark debuff's +12% may have been wrongly set as stackable,
   letting 3x archer_07 pile it up to +36%.
   Verify: force Mark max_stack to 1 and re-sim.

3. In 5:5, a 3-ranged + 2-chaser composition may simply be a setup that locks
   out enemy entry by range advantage (no synergy involved; the combo itself is strong).
   Verify: kill the Mark debuff effect to 0% and re-sim the same composition.

4. Given match_duration is longer than average (41s), team A may not be winning
   fast but rather "holding on without losing, then flipping it late."
   In that case the cause may lie in late-game scaling values, not win rate.
   Verify: extract and compare only the cumulative-damage curve at the 60s mark.

5. (Low probability) Sim seed bias — team A's first strike may have repeated
   only under the S0417 seed. Possibly a statistical artifact.
   Verify: re-sim the same composition across 50 seeds and check win-rate variance.

Stopping here is the point. The fifth hypothesis left a note of self-doubt — "this might just be seed bias" — and that is what good output looks like. If the LLM had instead declared "the cause is a stacking bug on the 표식 (Mark) debuff," that output itself would deserve suspicion. In balance work, the LLM's job is not diagnosis; it is narrowing the search space.

Step 3 — Run the Change Candidates Through the Sim (in Parallel)

Each of the five hypotheses carries a one-line verification sim. Instead of a human running them one by one, you bundle the change candidates and throw them in parallel. The core function, simulate_dps, looks like this in runnable form — the skeleton of that 18-year-old deterministic function.

def simulate_dps(attacker, target, formula, ticks=600, seed=0):
    """Deterministically sim one pair's combat. Same (input, seed) → same output."""
    rng = Rng(seed)                     # fixed seed → reproducible
    hp = target.hp
    total_damage = 0.0
    for t in range(ticks):              # assume 1 tick = 0.1s
        # defense factor: deterministic formula (the LLM does not write this)
        def_factor = target.defense / (target.defense + formula.def_const)
        raw = attacker.atk * (1 - def_factor)
        # crit: seed-based → same seed, same crit timing
        if rng.roll() < attacker.crit_rate:
            raw *= attacker.crit_mult
        # debuffs (Mark etc.) are injected deterministically by formula
        raw *= formula.debuff_multiplier(attacker, target, t)
        hp -= raw
        total_damage += raw
        if hp <= 0:
            return {"ttk": t * 0.1, "dps": total_damage / ((t + 1) * 0.1)}
    return {"ttk": None, "dps": total_damage / (ticks * 0.1)}  # not killed in time


def run_candidates(base_scenario, candidates, seeds=range(50)):
    """Sim each hypothesis's change candidate across 50 seeds in parallel. Also collect winrate variance."""
    out = {}
    for name, patch in candidates.items():           # patch = overrides part of formula
        scen = base_scenario.with_patch(patch)
        wins = [simulate_match(scen, formula=scen.formula, seed=s) for s in seeds]
        out[name] = {
            "winrate": mean(w["team_a_won"] for w in wins),
            "winrate_std": pstdev(w["team_a_won"] for w in wins),  # for verifying hypothesis 5
        }
    return out

Move the hypotheses into a candidates dictionary and run them all at once.

candidates = {
    "baseline(no_change)":   {},
    "hypo1_swap_hybrid":     {"team_a[3:5]": "refgame_melee_03"},
    "hypo2_Mark_max_stack1": {"skill.표식.max_stack": 1},
    "hypo3_Mark_effect0":    {"skill.표식.debuff": 0.0},
    "hypo5_seed_variance":   {},  # same composition, just 50 seeds
}
result = run_candidates(scenario_S0417, candidates, seeds=range(50))

The results (output in its actual execution form):

baseline(no_change)   winrate=0.91  std=0.04   ← not seed bias (hypothesis 5 rejected)
hypo1_swap_hybrid     winrate=0.74  std=0.06
hypo2_Mark_max_stack1 winrate=0.63  std=0.05   ← largest drop
hypo3_Mark_effect0    winrate=0.55  std=0.05   ← back near the average

The reading order is the diagnosis. Re-run the baseline across 50 seeds and the win rate is still 0.91 with a standard deviation of 0.04 — hypothesis 5 (seed bias) is rejected. Kill the Mark effect down to 0 and the win rate lands at 0.55, back near the average — the cause is indeed in the Mark debuff family. And forcing max_stack to 1 produced the largest drop, down to 0.63, so the heart of it is hypothesis 2 — the Mark debuff was stacking, and three archer_07 units piled it up to +36%. Of the five candidates the LLM threw out, no human had to verify all five; statistics settled it after running three.

Step 4 — A Human Adopts, and Leaves the Decision on Record

What the LLM did here was not say "Mark stacking is a bug." It merely put that hypothesis on the candidate list. Adoption is done by the balancer who reads the sim results — "Fix Mark's max_stack at 1. The all-archer_07 composition's win rate is still 0.63, above the 0.52 average, so in the next build, further adjust the Mark debuff value from 12% to 9% and re-measure."

That decision was made by a human, and its rationale — z=4.1 detection → 5 hypotheses → 3 sims → hypothesis 2 confirmed — fits in one line. The deterministic core stayed code to the end, and all the LLM did was swap a 40-minute wander for five lines of hypotheses. It never put a foot inside the core.

8.4.3 Five Slots, and the Cycle

The worked transcript above actually stepped through three of the five slots at once (anomaly detection, change exploration, anomaly interpretation). Unfolded into a cycle, the five slots turn like this.

flowchart TD
    A[Scenario definition] -->|Slot 1: auto scenario generation| B[Stat input]
    B -->|Slot 2: change candidate search| C{Deterministic simulation
simulate_dps} C --> D[Raw result JSON] D -->|find_outliers z-score| E[Anomaly pattern detection] E -->|Slot 4: LLM hypotheses 3–5| F[Hypotheses + verification sims] F -->|Slot 3: natural-language report| G[Balancer review] G -->|Slot 5: next-action proposal| H{Adopt / Reject} H -->|Adopt| B H -->|Reject| A style C fill:#dbeafe,stroke:#2563eb,stroke-width:2px style E fill:#dbeafe,stroke:#2563eb

Only the two blue nodes (the simulation and the z-score detection) are deterministic. The labels riding the remaining arrows — Slots 1 through 5 — are where AI attaches. Each time the cycle completes, the adopted change goes back into the stat input and the next sim runs. Turn this loop by hand and one revolution takes a day; with AI assistance, a few hours.

A quick pass over each of the five slots.

Slot 1 — automated scenario generation. Give it a one-line concept — "3:3 capture match, hold 3 flags for 1 minute to win, 10-second respawn" — plus one or two existing scenario yaml files, and the LLM fills out a new scenario yaml in the same schema. The balancer reviews only one thing: whether it slipped in rules that were never in the concept. The 1–2 hours of writing yaml from a blank page shrinks to a 15-minute review.

Slot 2 — change candidate exploration. The candidates dictionary in the worked transcript above is exactly this. Ask "what do we touch to raise tank survivability by +49%?" and the LLM throws out five candidates (base_def +50, a def_const adjustment, and so on); all of them go through the sim, and the one with the smallest side effects gets picked. Candidates are hypotheses; adoption belongs to the sim. This is the slot to treat with the most care — a bad candidate eats verification time.

Slot 3 — natural-language reports. Pull the metrics out of the sim's raw JSON with a script (deterministic), then hand only those metrics plus the change context to the LLM and have it write "the one page you take to the meeting": 3–5 lines of key changes, the top 5 affected characters, 2–3 follow-up actions. Nail down that it may not write any number beyond the metrics provided. Thirty minutes of raw cleanup becomes a 5-minute review.

Slot 4 — anomaly pattern interpretation. Steps 2–3 above are exactly this. The LLM attaches 3–5 hypotheses to the outliers the z-score picked out. The ban on definitive diagnoses is this slot's lifeline.

Slot 5 — next-action proposals. When the analysis ends, turn it into a prioritized checklist — "act now in this build / monitor for 1 week / re-review candidates in 1 week." It is a safety net that keeps the balancer from dropping a decision; it does not make the decision itself.

8.4.4 Where to Start, and How Far to Go

Turning all five slots on at once is the most common failure. Start from the exit side, where the payoff is large and the risk is small.

ROI ↔ adoption risk matrix (upper right = first) Higher ROI → Lower risk ↑ Slot 3 Reports ① Slot 4 Anomalies ② Slot 1 Scenarios ③ Slot 5 Actions ④ Slot 2 Changes ⑤ Most carefully, last

The circled numbers (①–⑤) are the adoption order. Slot 3 (reports) and Slot 4 (anomaly interpretation) sit in the upper right — high ROI (return on investment), low risk — so they go on first. With just these two running, throughput rises 2–3x, and more than 70% of the adoption payoff is recovered right here. Slot 2 (change proposals) is the red spot at the lower right; a bad candidate can eat verification time, so it goes on last, with the most care. Not every team needs all five, either — Slots 3 and 4 alone change a solo balancer's day.

A realistic sense of the rollout timeline (author's estimate, unverified — varies widely with team size and tool maturity): 1–2 weeks for Slot 3, 2 more weeks to add Slot 4, a month to add Slot 1, 2 weeks to add Slot 5, and Slot 2 last, at 1–2 months. It is another way of saying don't turn everything on at once.

8.4.5 Results, Costs, and the Most Common Traps

Here is what changed on my Project A after turning the five slots on over six months. The absolute figures are the author's estimates (unverified); trust only the direction and the ratios — the multipliers vary widely by environment.

Item Before After (direction)
Weekly sim cycles per balancer 5–7 25–35 (about 5x)
Report writing (each) 30–40 min 5-min review
Scenario writing (each) 1–2 hours 15-min review
Outlier found → diagnosed 1–2 days 4–6 hours
Measurement → next change decision 2–3 days 1 day

What matters here is not the multiplier but where the time went. Human time moved from cleaning raw data to making decisions. The number of balancers did not shrink; the territory one person can cover grew. Read the 5x throughput as a headcount cut and the whole point of the adoption drifts somewhere it should not.

The cost is small. With prompt caching applied, the monthly LLM bill for all five slots runs around $75 (author's estimate) — no more than 1/100 of one balancer's salary. So the real decision variable for adoption is not LLM cost but review burden: can you secure the time for a human to read and filter the hypotheses and reports the AI produces? That is the criterion for turning a slot on or off.

Finally, a few traps that have recurred in the same spots for 18 years, each with its remedy.

AI's place in balance work is clear: outside the deterministic core, in the five slots where humans used to wander. Keep the core as code to the very end and lift only the hand labor around it — that is how an 18-year-old simulator survives into the AI era.


Key Takeaways

One-Line Try It Yourself (Solo Scale-Down)

Next Chapter Preview

8.5 PvP and Competitive Balance — Win-Rate Matrix, Matchmaking, and Server Authority

The four chapters of this part so far have all fought a single enemy. How many seconds it takes to kill one boss, whether the tank survives 89% of the time, whether gold is leaking. All of it was a story of damage, survival, and income against a single target. In PvP, though, the enemy is a person. People don't move in fixed patterns the way a boss does, two players on the same class have different hands, and above all, they aim at each other's weaknesses. This is why it's common for a game's PvE balance to run deep while its PvP sits completely empty. Chapters 8.1 through 8.4 took the single-target DPS (damage-per-second) curve all the way to the end, but the web of counters — "scissors beats paper" — hasn't been drawn even once.

This chapter fills that gap. It covers three things — the win-rate matrix that captures matchups between classes and compositions, matchmaking and MMR, which decide who gets paired with whom, and server authority and anti-cheat, which can turn every one of those numbers into a lie. And the boundary that has run through this entire part holds here unchanged: combat formulas are deterministic; matchmaking and matchup detection get AI assistance. Not one step out of line.


8.5.1 The One Thing That Makes PvP Different from PvE

In PvE, a character's strength is an absolute value. If the swordsman's DPS is 800, it's 800, and the boss simply takes that 800. In PvP, strength is relative. The swordsman's 800 is enough against an archer, but against a shield bearer who reduces incoming damage by 30%, it gets cut to 560 and may fall short. The same character's strength changes depending on who the opponent is. This one fact makes PvP balance a fundamentally different problem from PvE.

So the unit of PvP balance is not one character's number but a pairwise relationship. The win rate of "swordsman vs. archer" and the win rate of "swordsman vs. shield bearer" each exist separately, and gathering all of these relationships produces a single table. The same class list runs along both axes, and each cell holds "the probability that the row beats the column." That is the win-rate matrix. Where PvE has the DPS curve, PvP has this matrix.

PvE is absolute, PvP is relational Swordsman DPS 800 Boss takes 800 as-is PvE: strength = absolute value Swordsman 800 Archer → 800 (effective) Shield bearer → 560 (short) PvP: strength = depends on the opponent → Win-rate matrix P(row beats column) Arch Shld Mage Sword .58 .42 .50 Arch -- .55 .47 Shld -- -- .61 (numbers are examples — not measured)

Reading the table on the right is simple. If the "swordsman vs. shield bearer" cell reads 0.42, the swordsman beats the shield bearer 42% of the time — the matchup favors the shield bearer. If every cell sits near 0.50, the balance is perfect, but a game like that is no fun. You need cyclical counters, rock-paper-scissors style, for class choice to mean anything. The problem comes when that cycle breaks somewhere and a cell appears in which one class beats everyone. If the tank at 2 a.m. was PvE's accident, "shield bearer above a 60% win rate against every class" is PvP's.

One thing I want to nail down up front. The numbers filling these cells (0.58, 0.42, and so on) are all examples, not measurements. Every game differs in class count, in skills, in target balance line. What to trust in this chapter is not the numbers but the structure — how the matrix gets filled, how it gets audited, and where AI attaches to that audit.


8.5.2 Simulation Fills the Matrix, AI Reads It

Filling one cell of the win-rate matrix uses exactly the same tool as the deterministic simulation in 8.4. Auto-simulate "swordsman vs. archer" 1,000 times, count how many the swordsman won, and that's the cell's win rate. With N classes there are N×N cells; run each cell 1,000 times and a table fills in. This simulation is code all the way down — given the same seed, it must reproduce the same matrix without a single character off. Only then is "the shield bearer got stronger in this build" not a lie.

There is one trap here that belongs to PvP alone. In a PvE sim, the enemy (the boss) follows a fixed pattern, but in a PvP sim the opponent has to choose actions too. You need a bot policy on both sides that decides how the swordsman fights. And if that bot is dumb, the entire matrix becomes a lie — pit two bots with terrible control against each other and you get a matrix where "the class that fires skills at random" wins, while in the hands of actually skilled players the result can be the exact opposite. So a PvP matrix must always carry a caveat: "what level of play does this bot imitate?" Bots are usually written as heuristics (use a skill the moment it's off cooldown, retreat below 30% HP, and so on), and the heuristic itself is deterministic.

Here is the skeleton of a bot policy in runnable form — a function that picks the same action for the same input, with no room for hallucination to creep in.

def bot_decide(me, enemy, cooldowns, t):
    """Deterministic bot policy. Same (state) -> same action. Not built by an LLM."""
    # 1) Survival first: evade/retreat if HP is below 30%
    if me.hp_ratio < 0.30 and cooldowns["escape"] <= 0:
        return Action("escape")
    # 2) Matchup skill: prioritize the mark unless the enemy is debuff-immune
    if cooldowns["mark"] <= 0 and not enemy.has("debuff_immune"):
        return Action("mark", target=enemy)
    # 3) Range management: open distance when a melee enemy closes in (ranged classes)
    if me.is_ranged and dist(me, enemy) < me.kite_range:
        return Action("reposition")
    # 4) Otherwise: the highest-damage skill that is off cooldown
    return best_ready_damage_skill(me, cooldowns)


def simulate_pvp_match(class_a, class_b, formula, seed=0):
    """Deterministically simulate one 1:1 match. Damage uses the formula from 8.1 as-is."""
    rng = Rng(seed)
    a, b = spawn(class_a), spawn(class_b)
    for t in range(MAX_TICKS):
        for me, foe in ((a, b), (b, a)):
            act = bot_decide(me, foe, me.cooldowns, t)
            apply_action(act, me, foe, formula, rng)   # formula = deterministic damage formula
        if a.hp <= 0 or b.hp <= 0:
            break
    return {"winner": "a" if b.hp <= 0 else "b" if a.hp <= 0 else "draw",
            "duration": t * TICK}

Filling an entire matrix is just the outer loop that runs this function 1,000 times per cell.

def build_winrate_matrix(classes, formula, n=1000):
    matrix = {}
    for ca in classes:
        for cb in classes:
            if ca == cb:
                continue
            wins = sum(
                simulate_pvp_match(ca, cb, formula, seed=s)["winner"] == "a"
                for s in range(n)
            )
            matrix[(ca, cb)] = wins / n          # rate at which ca beat cb
    return matrix

Up to here is the core, code to the end. AI attaches not to building this table but to reading it. With N = 8 there are 56 cells, and a human scanning 56 win rates by eye for "where is it broken" is the same labor as the 4 MB JSON at 2 a.m. Picking out the odd cells is exactly what 8.4's z-score detection already does.

def find_broken_cells(matrix, low=0.40, high=0.60):
    """Deterministically shortlist cells that deviate far from the balance line (0.5)."""
    broken = []
    for (ca, cb), wr in matrix.items():
        if wr > high or wr < low:
            broken.append((ca, cb, round(wr, 2)))
    return sorted(broken, key=lambda x: abs(x[2] - 0.5), reverse=True)

Once detection narrows the field, you hand that cell to the LLM. Same discipline as 8.4 — no definitive diagnoses; hypotheses and verification sims only. For example, give it the single line "shield bearer vs. mage 0.68 (largest z in the matrix)" and ask for 3\~5 possible causes, each with a one-line verification sim, like this.

[Broken cell]
Shield bearer → mage win rate 0.68 (balance line 0.50, largest z in the matrix)
Side data: average duration of this match 38s (overall average 22s)

[Related info]
- Shield bearer: -30% damage taken passive "Iron Wall", silence skill "Shield Bash" (2s)
- Mage: 70% of all damage concentrated in a skill with a 1.5s cast
- This matchup's frequency ranks high in the measured queue (popular pairing)

Request: 3~5 hypotheses for possible causes of this matchup collapse + 1 verification sim line each.
No definitive diagnoses. Stay at the level of "it may be...".

The LLM only throws hypotheses that narrow the search space, along the lines of "the -30% from the Iron Wall passive (철벽) and the 2-second silence may overlap into a positive feedback loop where the mage dies without ever landing its core cast / verify: cut the silence duration to 1 second and re-sim the same cell." Which one is actually true gets decided by running build_winrate_matrix again per candidate. That the match duration is 1.7 times the average — the LLM weaving that clue into the hypotheses, the kind of connection a human scanning 56 cells would easily miss, is the time AI earns at this spot.


8.5.3 Matchmaking: The Other Balance That Decides the Matchup

Even with a perfectly tuned win-rate matrix, the real reason a player feels "I lost" lies somewhere else: who they got matched against. When a player at skill 1500 meets one at 2200, the result is already decided even if the class matchup is 5:5. So matchmaking is not a mere server feature — it is part of balance. If the matrix is in charge of fairness between classes, matchmaking is in charge of fairness between skill levels.

Most competitive games keep an MMR (Matchmaking Rating). It is a hidden score that rises when you win and falls when you lose, and players with similar scores get paired. The score update is a deterministic formula — the Elo rating system is the most widely used, and being a public standard, it is one of the few formulas this book is allowed to quote.

# Elo: public standard update formula (not a made-up value)
expected_a = 1 / (1 + 10 ** ((rating_b - rating_a) / 400))
new_rating_a = rating_a + K * (score_a - expected_a)
#   score_a: 1 on a win, 0 on a loss
#   K: update strength constant (set by the game; usually chosen in the 16~40 range)
#   400, 10: constants fixed in the Elo definition

The formula itself is deterministic, and it is no place for AI. But matchmaking carries one tension that deterministic formulas alone cannot resolve: the fairness ↔ wait-time trade-off. Pair only opponents with the exact same score and the match is fair, but when no such opponent is in the queue, the player waits 10 minutes. Allow a generous score gap and matches come fast but unfair. The tension sharpens in late-night hours, on unpopular classes, and in high-rating brackets.

flowchart LR
    A["Match request
user MMR 1500"] --> B{"Opponent within ±50 in queue?"} B -->|Yes| C["Match immediately
fairest"] B -->|No| D["Widen tolerance as wait grows
±50→±200"] D --> E{"Match made?"} E -->|Made| F["Match
fairness ↕ wait ↕ balanced"] E -->|Timeout| G["Insert bot / hold the queue
policy decision"] C --> H["Elo update (deterministic)"] F --> H style H fill:#dbeafe,stroke:#2563eb,stroke-width:2px style D fill:#ffedd5,stroke:#ea580c style G fill:#ffedd5,stroke:#ea580c

Only the blue node (the Elo update) is deterministic. The orange nodes — when and by how much to widen the tolerance, what to do on timeout — are where AI assistance reaches. Even here, though, AI does not make real-time matching decisions. That is server logic that must be fast and reproducible, so it belongs to rule-based code. What AI attaches to is the analysis used to tune those rules: summarizing "in last week's matchmaking logs, which rating brackets, time slots, and classes had poor match quality (win-rate skew, wait times)," and proposing candidates for "which segment's wait time shrinks if the tolerance curve changes this way." It is 8.4's position 3 (reports), position 4 (anomaly interpretation), and position 2 (exploring change candidates), with the stage moved to matchmaking logs.

It's worth marking where matchmaking entangles with the win-rate matrix. If the matching algorithm only equalizes scores and ignores class, broken matchup cells get exposed as-is. If the cell where the shield bearer beats the mage 68% of the time is still alive and matchmaking keeps pairing the two, mage players stack up perceived defeats even heavier than the matrix number suggests. So the matrix audit and the matchmaking-log analysis do not run separately — they are the entrance and the exit of the same cycle: fix the broken cell in the matrix, then check in the matchmaking logs how often that cell actually got paired.


8.5.4 Server Authority: The Precondition for Balance

Everything so far — the matrix, MMR, the sims — has quietly rested on one assumption: that the results clients report are true. In PvE this is rarely a problem. You're killing a boss alone; who would you cheat? But in PvP there is an opponent, winning raises your score, and so a motive to cheat exists. The moment a client shows up that forges damage, forges position, and ignores cooldowns, the deterministic formula of 8.1 is deterministic only on paper. On the live server, someone's swordsman is dealing twice what the formula says.

So the first balance rule of a competitive game comes before the matrix: never let the client decide outcomes. Damage calculation, cooldown checks, hit checks — every computation that touches balance is the server's authority. The client sends inputs only (move where, use which skill), and whether that input fits the formula, whether the cooldown has elapsed, whether the target is in range — the server re-verifies all of it. A client-sent "damage 999" gets ignored; only the value the server computes from the formula is applied.

Server authority = the sole enforcer of the balance formula Client sends inputs only "use skill 1, coords (x,y)" cannot decide outcomes input verified result Server (authority) verifies cooldown/range damage = formula (8.1) deterministic · enforcement ignores "damage 999" Anomaly logs impossible inputs pattern detection AI assist allowed

When server authority collapses, the entire balance effort becomes a lie. However precisely you tune the win-rate matrix, if one class is forging damage on live, that matrix is a promise on paper. So anti-cheat is not a separate security chore — it is a question of how trustworthy your balance data is. When live win rates diverge sharply from the simulated matrix, the first thing to suspect should be not "is the formula wrong" but "is this data clean."

Here AI's place becomes clear once again. The cheat verdict itself — "this input is void" — is the work of deterministic rules. An input that moved 30 meters in 0.1 seconds is physically impossible, so a rule blocks it. The same input must yield the same verdict, and you cannot afford wrongful suspensions, so a probabilistic LLM cannot sit there. On the other hand, shortlisting anomalous patterns as candidates is within AI assistance's reach: combing server logs for candidates like "this account's hit-rate distribution sits z-something away from the human distribution" or "this group of accounts shares the same abnormal pattern," and putting them up for human review. Port the table from 8.1 over to PvP, and the boundary looks like this.

Area AI Why
Server damage, hit, and cooldown verdicts Never Deterministic core. If same input = same verdict breaks, fairness collapses
Elo/MMR score updates Never Public-standard deterministic formula. Shake it and rankings become lies
Cheat-ban verdicts themselves Never No wrongful suspensions allowed. Same evidence = same verdict
Win-rate matrix simulation Never If it can't reproduce, "this class got stronger" becomes a lie
Detecting and interpreting broken matchup cells Allowed z-score shortlists the cells, LLM hypothesizes (no definitive diagnoses)
Matchmaking-log quality analysis and tuning candidates Allowed Proposes change candidates for the wait/fairness trade-off (sim-verified)
Extracting suspected-cheat pattern candidates Allowed Candidates for human review only. Ban decisions are human + rules

The line is identical to 8.1, word for word. AI lives only outside the deterministic core. The enforcing inside — damage, scores, bans — is the rulebook; the outside, which detects, interprets, and pushes candidates, is AI's place.


8.5.5 Tying It into One Cycle

The three topics — matrix, matchmaking, server authority — are not three jobs running separately. They are three segments of a single competitive-balance cycle. Server authority guarantees clean data, that data drives the matrix audit, the matchmaking logs confirm how the audit's results actually land in the live queue, and sims verify candidates again before they go into the build.

flowchart TD
    A["Server authority + anti-cheat
clean live data"] --> B["Live win-rate matrix
(measured)"] B --> C{"Diverges sharply from
the simulated matrix?"} C -->|"Diverges → suspect the data"| A C -->|"Matches → balance problem"| D["Broken-cell detection z-score"] D -->|"LLM hypotheses 3~5"| E["Change candidates + verification sims"] E --> F["build_winrate_matrix
re-sim per candidate (deterministic)"] F --> G["Balance designer adopts/rejects"] G --> H["Apply to build (irreversible)"] H --> I["Matchmaking log analysis
perceived-quality check + tuning candidates"] I --> A style A fill:#dbeafe,stroke:#2563eb,stroke-width:2px style F fill:#dbeafe,stroke:#2563eb,stroke-width:2px style D fill:#ffedd5,stroke:#ea580c style E fill:#ffedd5,stroke:#ea580c style I fill:#ffedd5,stroke:#ea580c

The blue nodes (server authority, sim recomputation) are deterministic; the orange nodes (detection, hypotheses, matchmaking analysis) are AI assistance. The most common failure in this cycle is skipping the C branch. When the live matrix diverges from the sim and you reach straight for the formula, you end up chasing cheat-polluted data and nerfing a perfectly healthy class. This one branch — suspecting data cleanliness first — plays the same role in PvP that 8.1's "change history" did. Skip it, and 2 a.m. comes back.

Finally, a few traps from 18 years around PvP balance, each with its prescription.

AI's place in PvP is the same as in PvE: detection, interpretation, and candidates outside the deterministic core — damage, scores, bans, sims. Guard the core with code and server authority, and shave off only the hand labor of a person scanning 56 cells and wandering through matchmaking logs — this chapter is the final proof that this whole part moves on one skeleton.


Try It Yourself — Auditing a Single Win-Rate Matrix

setup. Extend 8.4's simulate_dps into a 1:1 simulate_pvp_match, plus a bot_decide (deterministic heuristics) to drive both bots. First confirm that a fixed seed reproduces the same matrix. Pull the damage formula straight from 8.1, unchanged, and record in one line what level of play the bot imitates.

prompt. Use find_broken_cells to shortlist the cells outside the balance band (0.40\~0.60), then hand only the one cell with the largest z to the LLM.

For the attached broken cell (shield bearer vs. mage 0.68, match duration 38s vs. average 22s),
propose 3~5 hypotheses for possible causes and 1 re-sim line to verify each.
Related skill/passive info is attached below. No definitive diagnoses — "it may be..." only.
Do not edit the numbers directly; propose candidates only.

verify. Don't take the AI's hypotheses on faith. Feed each hypothesis's change candidate into build_winrate_matrix, re-sim with the same seed, and check both that the cell returns toward 0.50 and that no other cell breaks (a PvP change tends to wreck the next cell over while fixing this one). Adopt only candidates that satisfy both conditions, and as in 8.1, leave the rationale, the rejected candidates, and the predicted values in the decision log. One week after the build ships, append the live measured win rate to that log.

Solo Scale-Down

Even a solo prototype with only two classes and no server keeps the same skeleton. A 2×2 matrix is enough, and the sim is 8.1's 30-line loop plus a single line of bot policy (use the biggest damage skill off cooldown). For server authority, just uphold the principle — "the client doesn't get to decide outcomes" — in your code structure; full anti-cheat isn't needed before you have users. Skip MMR at first too, and just run 1,000 matches to check whether the matrix tilts past 60% to one side. Use AI only to read the result and summarize "which matchup broke and why it might have." The one line to hold at any scale: damage and win/loss are decided by code and the server — never by the LLM.


Key Takeaways

Next Chapter Preview

Part 9 · Ux Ui Design

9.1 Running the HUD Screenshot Through Lint — Where AI Catches Out-of-Gaze Placement and Failing Contrast

Primary reader: the UX designer responsible for HUD and UI (on a mid-size team of 10–50) Condensed version for solo/hobbyist readers: §9.1.8, "If You're Solo, Just This Much"

The day we put a new debuff alert on the HUD in a QA build, the designer said it "read fine." The next day, the player forums said "I died because I couldn't see the debuff." The alert sat in the center of the screen, pale yellow text on a gray background. It was visible on the designer's monitor — and invisible on a 6-inch phone with combat explosion effects covering the screen. The real problem was that this wasn't the first time. Every build, every screen, the same kind of accident repeated under "it'll probably be fine this time."

This chapter focuses on the one piece of work that breaks that cycle: a lint gate — a quality gate, in CI terms — that takes a single finished HUD screenshot as input and automatically detects whether any P0 element has drifted out of the zones the eye actually reaches (the top status band and the two bottom action corners), and whether text contrast clears the readability threshold. The general principles of HUD design — priority tables, gaze flow, platform splits — are already well covered in other books, so this chapter spends its pages only on the review loop that enforces those principles automatically on every build. The point is to make AI look at a screen and say, in coordinates and numbers, "this text is at 2.0:1 contrast, which fails WCAG 4.5:1." It replaces the "it looks fine to me" argument with code and standards.


9.1.1 Review Criteria Are Published Standards, Not Feelings

HUD reviews reach a different conclusion with every reviewer because the criterion is subjective: "it reads fine" versus "it doesn't." Fortunately, much of readability and accessibility has already been nailed down in numbers by standards bodies. There is nothing to make up.

Review item Standard threshold (source) Automatable verdict
Body text contrast 4.5:1 or higher (WCAG 2.1 SC 1.4.3) Yes — computed from foreground/background color values
Large text (18pt+) contrast 3:1 or higher (WCAG 2.1 SC 1.4.3) Yes
Non-text (icons, gauges) contrast 3:1 or higher (WCAG 2.1 SC 1.4.11) Yes
Minimum touch target size 44×44 pt (Apple HIG) / 48×48 dp (Material) Yes — from element size
Thumb reach zones Under a two-handed landscape grip, the bottom-left and bottom-right corners rate "easy" (left thumb = movement, right thumb = skills). The industry-standard thumb-zone model Partial — via zone rules

Only the last row (thumb reach) is an industry convention rather than a quantitative pass line; the four rows above it are pass lines published by W3C, Apple, and Google. Contrast is the clearest of all. WCAG even publishes the formula: take the relative luminance of the two colors and compute (L1+0.05)/(L2+0.05). Put pale yellow (#D4C84A) text on a gray (#888) background into that formula and you get about 2.0:1 — below 4.5:1, an unambiguous fail by the standard. "It was visible on the designer's monitor" is not a rebuttal that survives here.

One thing needs to be stated plainly. For MMORPGs and RPGs, the mobile screen standard is landscape. The reasons are information density and controls. At the same screen size, a landscape grip fits more always-on information per screen than portrait, and both thumbs can operate the left side (movement) and the right side (skills) simultaneously. A one-handed portrait grip suits casual puzzle and idle games; it does not suit an MMORPG with heavy concurrent information and two-handed controls. So every gaze and placement verdict in this chapter assumes a two-handed landscape grip. The screen divides into a horizontal status band along the top, two bottom action corners on the left and right, the central game area between them, and a bottom-center slot bar below the game area (consumables, auto-use items, quick slots).

These five rows are the review rulebook this chapter hands to AI. Only when you can say "the debuff text is at 2.0:1 contrast, violating SC 1.4.3" — not "the debuff seems kind of hard to see" — do human review and AI review produce the same verdict.

Putting the platform baselines side by side with PC clarifies the starting point of the review. Project A is mobile-first with PC as the secondary platform, so the rulebook carries both.

Criterion PC (secondary platform) Mobile (primary platform, landscape)
Screen and input 27-inch+ / 1px mouse precision, hover, hotkeys 6.x-inch landscape / two thumbs, no hover
Concurrent always-on info Handles 30–50 kinds 12–16 kinds is the limit (author's estimate, unverified)
Gaze and control reach The entire screen (the cursor reaches anywhere) Only the top status band, the bottom-left/right corners, and the bottom-center slot bar are "easy"
Precision 1px clicks 44pt minimum touch target (HIG)
Main review risk Cognitive load from information overload Small screen + finger occlusion + center burial

On PC, mouse precision, hover tooltips, and a large display mean gaze and controls reach even a dense screen. Mobile in landscape beats portrait but still holds less than PC; pressable elements are tied to the two thumb corners, and with no hover, P0 information must stay permanently visible. So the essence of mobile HUD review is not "does it look nice" but "is every P0 element where the eye reaches (the top band or the two corners), and does the text clear the standard contrast ratio?" Nailing that verdict to a standard, so it doesn't wobble from person to person, is this chapter's job.


9.1.2 [Worked Transcript] Running One HUD Screenshot Through Lint

Here is one full cycle of how this actually runs. What follows is a faithful reproduction of a combat HUD review session from my project (a mobile-first MMORPG, "Project A" from here on). The input prompts can be copied as-is; the outputs are reconstructions of the real session.

Step 1 — Input: Throw the Screenshot and the Element Manifest Together

Throw in only a screenshot and AI "guesses" at the screen. So I attach a manifest of what the build already knows — per-element coordinates, colors, and classifications. This is not something you write fresh; you only extract it from build artifacts (the realities of extraction are compared honestly in §9.1.4).

# hud_capture_manifest.yaml — bundled with the QA build screenshot
screen: { w_pt: 844, h_pt: 390 }   # 6.x-inch landscape, in pt (landscape grip)
elements:
  - id: hp_bar        # HP bar
    class: P0
    rect_pt: [12, 18, 150, 16]      # x, y, w, h — top left
    fg: "#FF5A5A"  ; bg: "#1A1A1A"
  - id: skill_slot_1  # skill slot (right thumb)
    class: P0
    rect_pt: [760, 300, 40, 40]     # ← bottom-right corner, note the size
    fg: "#FFFFFF"  ; bg: "#202830"
  - id: debuff_alert  # debuff alert (added yesterday)
    class: P0
    rect_pt: [400, 180, 70, 24]     # ← screen center, note the position
    fg: "#D4C84A"  ; bg: "#888888"   # ← note the contrast
  - id: minimap
    class: P1
    rect_pt: [744, 20, 80, 80]       # top right
    fg: "#A0C0FF"  ; bg: "#101820"

(Korean comments in the manifest: the file header reads "bundled with the QA build screenshot"; 체력바 = HP bar; 스킬 슬롯 (우엄지) = skill slot, right thumb; 디버프 알림 (어제 추가) = debuff alert, added yesterday; 우상단 = top right. The arrow notes flag what to watch: the corner slot's size, the alert's center position, and its contrast.)

Step 2 — The Prompt: Ask for a Review, but Force the Standard and the Format

The attached screenshot is Project A's combat HUD (two-handed landscape grip), and the yaml holds per-element coordinates, colors, and classes for that screen. Cross-check the two and review.
Compute WCAG contrast from fg/bg and write out the numbers — FAIL below 4.5:1 for text, 3:1 for icons and large text.
WARN if a P0 element leaves the top status band or the bottom-left/right corners and floats in the screen center (the center gets buried under combat effects).
FAIL any interactive element under 44pt, or outside the thumb corners and the bottom-center slot bar.
Report anything visible on screen but missing from the manifest, and pass anything you're not sure about to me marked "ambiguous."
Table only (element | check | measured value | verdict | notes), no prose.
// (intent: P0 = information that must be found within 0.5 seconds / consumables, auto-use, and quick slots belong at the bottom center)

(The prompt, in brief: the attached screenshot is Project A's combat HUD — two-handed landscape grip — and the YAML holds per-element coordinates, colors, and classes; cross-check the two. Compute WCAG contrast from fg/bg and show the numbers — FAIL below 4.5:1 for text, 3:1 for icons and large text. WARN if a P0 element leaves the top status band or the bottom corners and floats in the screen center (the center gets buried under combat effects). FAIL any interactive element under 44pt or outside the thumb corners and the bottom-center slot bar. Report anything visible on screen but missing from the manifest, and pass anything uncertain to me marked "ambiguous." Table only — element | check | measured value | verdict | notes — no prose. The trailing comment states the intent: P0 = information that must be found within 0.5 seconds; consumables, auto-use, and quick slots belong at the bottom center.)

Note that more lines went into how to output the results than into the review rules themselves. Give AI free rein and it produces compliments like "overall this looks clean." So I nail down five things in advance: the calculation (WCAG contrast), the pass lines (4.5:1 and 44pt), the output format (a table), the obligation to show measured values, and an escape hatch — "if you're not sure, hand it to a human." Those five nails turn compliments into measurements.

Step 3 — Claude's Output (Reconstructed, Raw)

| element     | check     | measured value            | verdict | result notes                          |
|-------------|-----------|---------------------------|---------|---------------------------------------|
| hp_bar      | contrast  | #FF5A5A/#1A1A1A = 5.7:1   | PASS    | non-text gauge, meets 3:1             |
| hp_bar      | placement | top left (12,18)          | PASS    | P0, top status band                   |
| skill_slot_1| touch     | 40×40 pt, bottom-right corner | FAIL | below HIG 44pt, short by 4pt (corner OK)|
| skill_slot_1| placement | bottom right (760,300)    | PASS    | right-thumb reach corner              |
| debuff_alert| contrast  | #D4C84A/#888888 = 2.0:1   | FAIL    | below 4.5:1 for body text (SC 1.4.3)  |
| debuff_alert| placement | screen center (400,180)   | WARN    | P0 yet centered — buried under combat effects |
| minimap     | contrast  | #A0C0FF/#101820 = 9.8:1   | PASS    |                                       |
| minimap     | placement | top right (744,20)        | PASS    | P1, allowed on right of top status band |

Additional report:
- The manifest lists 4 elements, but one more blinking yellow icon shows
  at the bottom left of the screenshot (estimated coords ~70,330). Suspected
  manifest omission. (Ambiguous — can't tell from the screen alone what it is)

Summary: 2 FAIL (skill_slot_1 touch, debuff_alert contrast), 1 WARN (debuff_alert
placement), 1 ambiguous (unregistered icon).

(Output columns: 요소 = element, 검사 = check, 측정값 = measured value, 판정 = verdict, 비고 = notes. The checks are 대비 = contrast, 배치 = placement, 터치 = touch. The "additional report" (추가 신고) at the bottom flags one extra blinking yellow icon at the bottom left — estimated coordinates \~70,330 — that is missing from the four-element manifest, marked "ambiguous (애매) — can't tell from the screen alone what it is." Summary line: 2 FAIL — skill_slot_1 touch, debuff_alert contrast; 1 WARN — debuff_alert placement; 1 ambiguous — the unregistered icon.)

The most valuable part of this output is not the pass/fail table but the "additional report" and the "ambiguous" entry at the very bottom. AI caught a blinking icon on screen that the manifest didn't list, and handed it to a human because it could not determine what it was. A good prompt makes it possible for AI to say "this one, I don't know."

Step 4 — Verification and Rejection (The Human's Place)

Do not accept this output as-is. The AI's review itself gets one round of human review. In this session, one verdict was in fact overturned by a human.

The contrast FAIL and placement WARN on debuff_alert are correct. Pale yellow on gray violates the standard exactly as §9.1.1 showed, and putting a P0 alert in the center of a landscape screen is the classic mistake of letting combat effects bury it. So far, AI was right.

The problem is the touch FAIL on skill_slot_1. AI took the manifest's 40×40 pt at face value and ruled "below 44pt" — but in the actual build, this slot draws at 40pt while its touch hitbox extends 6pt on every side, making the real tap area 52pt. The manifest's rect_pt carried only the drawn rectangle, not the hitbox — a defect in the input data, not a misjudgment by AI. AI judged precisely within the data it was given (the corner-position verdict was correct), and the human knew a build fact the code didn't: the hitbox expansion. This FAIL, the human strikes down.

So two things happen at once: fix the manifest extraction script so it dumps hitboxes too (repairing the data defect), and re-request from AI.

skill_slot_1 draws at 40pt visually, but its hitbox extends 6pt on every side, so the real tap area is 52pt (hit_rect added to the manifest). Re-check touch against this.
Leave the debuff_alert FAIL/WARN as-is, and propose 3 color combinations that clear 4.5:1 contrast (keep the yellow family, darken the background). Also give one coordinate to move it from the center to the right side of the top status band.

AI corrected skill_slot_1 to PASS against the 52pt hitbox, returned three color combinations for the debuff that darken the background to #2A2A00 and reach 7.8:1, and gave one coordinate that moves the alert to the right side of the top status band (around 600,18). One round trip, done. Sweep the screen by eye every build and the same accidents repeat; run the screenshot plus manifest through lint and the contrast, placement, and touch violations drop out as numbers, leaving humans to judge only the exceptions code doesn't know (hitboxes) and the ambiguous cases (unregistered icons) (one screen takes a dozen-plus minutes by hand, a few minutes with this loop — author's estimate, an unverified hypothesis; read it less as absolute times and more as the structural difference between "sweeping by eye" and "measuring against a standard").


9.1.3 Landscape HUD Gaze and Placement — Why the Center Is Dangerous

Capture in a single diagram why debuff_alert earned its WARN in the session above and where P0 information belongs, and every placement verdict afterward gets faster. On a phone held in landscape, the screen divides into four places: the horizontal status band at the top (where the gaze lands first and fingers never go — read-only), the two bottom corners left and right (where both thumbs rest — left thumb = movement, right thumb = skills), the central game area between them (where combat happens), and the bottom-center slot bar below the game area (consumables, auto-use items, quick slots and skill slots). In the figure below, green and amber are where P0 elements and slots are safe; red is the game center where P0 alerts get buried.

Top horizontal status band — first gaze priority (HP · MP · target, read-only) Center — game area (effect overload) P0 alerts get buried — where debuff_alert was flagged Bottom center — consumables · quick slots · auto Potion Auto Slot Left thumb Move Right thumb Skills HP MP Target Map P1 Debuff? Move Skill Skill Skill

The rule is simple. P0 information (HP, MP, critical alerts) goes inside the green — the top horizontal status band or the two bottom corners. Those are the paths the gaze hits first or where the thumbs always rest. Conversely, the game center (red) is where combat itself happens — put a P0 alert there and the moment effects flood the screen, the information is buried. One caution: the game center and the bottom center are not the same. The game center is dangerous, but the bottom-center slot bar (amber) below it is where consumables, auto-use items, quick slots, and skill slots live. It sits between the two thumbs so that what I use, or what gets consumed automatically, stays in view at a glance. And the mapping is threefold: read-only information (HP/MP/target health) at the top, pressable elements (movement, skills) in the two bottom corners, consumables and slots at the bottom center — these are the finger and gaze zones. This one diagram explains why the debuff alert took a WARN in §9.1.2: a P0 element that must be seen within 0.5 seconds was placed, of all places, in the game center where it is least visible. The fix that moved it to the right side of the top status band sent it straight back into this diagram's green.


9.1.4 How to Extract the Coordinates — Implementation Honesty

This chapter's lint stands on the premise that clean per-element coordinates and colors come in. But where and how you extract those coordinates is, in practice, the most consequential fork in the road. It's the part books usually fudge, so here is an honest comparison of the three paths. There is no single right answer; it splits by team situation.

Path What it does Strength Weakness / reality
① In-game telemetry log The build itself dumps the coordinates, sizes, and colors of the widgets the UI framework draws Coordinates are exact (not estimates); hitboxes and anchors come out too A dump hook must be planted in UI code. Requires programmer collaboration. Once in place, the most trustworthy
② Off-the-shelf vision API Feed the screenshot to an OCR/object-detection API and extract text and box coordinates No build changes needed; works on external screenshots too Coordinates are approximations; weak at classifying non-text such as gauges and icons. Sending data out = leak risk for an unreleased build
③ Build it yourself (pixel analysis) Read the screenshot directly and extract color boundaries and boxes heuristically Minimal dependencies; sufficient for contrast math Knows nothing of element meaning (is this P0?). Useful only when cross-checked against a manifest. Maintenance burden

The relationship among the three paths explains this chapter's worked transcript exactly. In §9.1.2, the contrast checks were accurate because the color values (fg/bg) came in exactly via ① or ③, and the touch FAIL was overturned by a human because the hitbox was missing from the manifest (② and ③ cannot see hitboxes; only ① can). In other words, contrast can be caught from pixels alone, but touch hitboxes cannot be caught without ① telemetry. Know this limit going in, and you know where to draw the line on how far to trust AI review results.

My project's choice is the structure of ① telemetry as the source of truth, with AI as the reviewer cross-checking the screenshot against the telemetry manifest. AI catches what shows on screen but isn't in the manifest (the unregistered blinking icon of §9.1.2); humans catch what's in the manifest but wrong on screen. Either one alone leaves blind spots on both sides.

flowchart LR
    A["QA build
HUD screenshot"] --> C B["telemetry dump
coordinates, colors, hitboxes
(path ①)"] --> C["manifest + screenshot"] C --> D["AI review
contrast, placement, touch
+ report unregistered elements"] D --> E{"WCAG/HIG
standard verdict"} E -->|FAIL/WARN| F["human review
exceptions and ambiguous only"] F -->|re-request| D F -->|pass| G["build gate passed"] classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; class A,B,C data; class D ai; class E code; class F human; class G pass;

Human hands touch only two places: feeding in a clean telemetry dump (the front), and judging the exceptions code and standards can't catch (hitboxes) and the ambiguous cases (unregistered elements) at the back. The tedious contrast math and placement cross-checks in between, AI and the standards grind through.


9.1.5 From Rulebook to Code — Automatic Gates for Contrast, Touch, and Corners

Having AI redo the arithmetic every time costs tokens and time. Items that resolve deterministically — contrast, touch size, corner reach — get hit by code first. AI steps in only for what code can't catch: interpreting what the screen means, spotting unregistered elements. The two are not in competition; they divide the work.

# hud_lint.py — standards validation for the HUD manifest (skeleton)
# input: telemetry manifest (per-element rect/hit_rect/fg/bg/class/interactive)
# output: list of WCAG/HIG + two-handed reach violations

def _luminance(hex_color):           # WCAG relative luminance
    r, g, b = (int(hex_color[i:i+2], 16) / 255 for i in (1, 3, 5))
    f = lambda c: c/12.92 if c <= 0.03928 else ((c+0.055)/1.055) ** 2.4
    R, G, B = f(r), f(g), f(b)
    return 0.2126*R + 0.7152*G + 0.0722*B

def contrast_ratio(fg, bg):          # WCAG contrast ratio
    L1, L2 = sorted((_luminance(fg), _luminance(bg)), reverse=True)
    return (L1 + 0.05) / (L2 + 0.05)

def in_thumb_corner(e, w, h):
    """Is this in a bottom-left/right corner a thumb reaches under a two-handed landscape grip?"""
    x, y = e["hit_rect"][0] / w, e["hit_rect"][1] / h
    bottom = y > 0.55
    left_corner  = bottom and x < 0.30   # left thumb = movement
    right_corner = bottom and x > 0.70   # right thumb = skills
    return left_corner or right_corner

def lint(elements, screen_w, screen_h):
    issues = []
    for e in elements:
        # rule A: contrast ratio (text 4.5:1 / non-text & large text 3:1)
        need = 4.5 if e["kind"] == "text" else 3.0
        cr = contrast_ratio(e["fg"], e["bg"])
        if cr < need:
            issues.append(f"[A] {e['id']}: contrast {cr:.1f}:1 < {need}:1 (WCAG SC 1.4.3)")
        # rule B: touch target — based on the hitbox (not the visual size)
        if e.get("interactive"):
            tap = min(e["hit_rect"][2], e["hit_rect"][3])   # ← hit_rect, not rect
            if tap < 44:
                issues.append(f"[B] {e['id']}: tap {tap}pt < 44pt (HIG)")
            # rule C: interactive elements must sit in the two thumb corners (bottom left/right)
            if not in_thumb_corner(e, screen_w, screen_h):
                issues.append(f"[C] {e['id']}: interactive element placed outside the two-handed thumb corners "
                              f"(x={e['hit_rect'][0]}, y={e['hit_rect'][1]})")
    return issues

(In the code's Korean strings and comments: rule A is contrast — 대비 = contrast; rule B is the touch target checked against the hitbox, not the visual size — 탭 = tap; rule C requires interactive elements to sit in the two thumb corners — the message reads "interactive element placed outside the two-handed thumb corners." The docstring of in_thumb_corner asks: is this in a bottom-left/right corner that a thumb reaches under a two-handed landscape grip?)

This code ends the meeting-room skirmish of "isn't this text a little hard to read?" When code prints [A] debuff_alert: 대비 2.0:1 < 4.5:1 (WCAG SC 1.4.3) ("대비" = contrast), there is nothing left to debate. You fix it. Two lines deserve attention: rule B reads hit_rect, not rect, and rule C passes interactive elements only in the two bottom corners, left and right — the lesson from §9.1.2 where a human overturned AI (the hitbox), and the reach limits of a two-handed landscape grip, both written into code. The key to the landscape verdict is that it checks the left-thumb corner (movement) and the right-thumb corner (skills) separately, rather than a single "thumb arc" threshold. An exception a human catches once, code catches from then on. That leaves AI a deliberately narrow role: "skip what code already passed; report only what shows up on screen alone — unregistered elements, visual overlap, clipping." What resolves deterministically goes to code, what needs on-screen interpretation goes to AI, and the exceptions that require knowing the build go to humans — that division is the heart of it.


9.1.6 Where This Chapter's Numbers Come From

The numbers in this chapter have only three kinds of sources. Contrast 4.5:1, touch 44pt, and 48dp are the official values of WCAG SC 1.4.3, the HIG, and Material; pale yellow #D4C84A on a #888 background coming out to about 2.0:1 is a computed value — those color values plugged into the published formula (§9.1.1, §9.1.5). "A dozen-plus minutes per screen by hand, a few minutes with the loop" and "12–16 kinds of always-on landscape information" are unverified author's estimates and are labeled as such in the body. The rest — contrast FAIL counts per build, undersized touch hitboxes, thumb-corner violations, telemetry mis-tap rates — are values you can count directly from build logs. Outcome metrics like player complaint counts, where causation can't be pinned on the HUD alone, were not promoted to KPIs.


9.1.7 Common Failures

Pattern Why it fails Prescription
Eyeball review on the designer's monitor Misses the 6-inch screen and combat-effect conditions, so contrast accidents repeat Make screenshot lint a build gate (§9.1.2)
Throwing only a screenshot at AI with "review this" It guesses coordinates and returns approximate verdicts — not trustworthy Attach the telemetry manifest (§9.1.4)
Judging touch targets by visual size Misses hitbox expansion and FAILs perfectly fine buttons Check against hit_rect (§9.1.5)
Placing P0 alerts in the screen center Buried under combat effects — "died because I couldn't see it" Move to the top status band or the two corners (§9.1.3)
Placing control buttons at the screen's mid-left or top Thumbs can't reach them in a two-handed landscape grip Move to the bottom-left/right corners (§9.1.5, rule C)
Designing around a one-handed portrait grip Two-handed landscape is the MMORPG standard; neither the information nor the controls fit Switch to two-handed landscape (§9.1.1)
Debating contrast as "I can see it / I can't" Conclusions differ from person to person Use the computed WCAG 4.5:1 value (§9.1.1)

The fourth one repeats most often. When a new alert gets added in a hurry, the only empty space is the screen center, so that's where it lands — and that center is exactly where the game happens.


9.1.8 Try It Yourself — One Step You Can Take Today

If you're solo, just this much: You don't need telemetry or a manifest. Take one landscape HUD screenshot of your own game (or a game you love), eyedropper the foreground/background colors of the two or three smallest pieces of text or icons, jot them down by hand, and run §9.1.2's prompt once. Pick one contrast figure AI computed and re-verify it yourself with an online WCAG contrast calculator — you'll feel firsthand how "I can see it / I can't" turns into a number. If AI lets a P0 element sit in the screen center, push back: "look again at why the center is dangerous."

If you're on a team, start with this one step. Agree with a programmer on the telemetry hook that dumps widget coordinates, colors, and hitboxes from the UI framework (path ①), and put just the contrast_ratio function from §9.1.5 into the build. The contrast formula is a published standard, so there's no arguing with it, and that one function alone makes contrast FAILs drop out as numbers on every build. Layer in_thumb_corner on next, and code starts catching control elements that stray from the corners under a two-handed landscape grip. Interpretation — placement, unregistered elements — goes on top of that, with AI.


Key Takeaways

Next Chapter Preview

9.2 Skill Button Layout — AI Drafts Three Layouts, lint Rejects Them

Primary audience: UX and combat designers on mobile-first action games and MMORPGs (mid-size teams) Scaled-down version for solo/hobbyist readers: §9.2.7 "Solo Scale-Down"

Where and how do you lay out a new class's six skills on a mobile screen? Whenever this question reached a meeting, the first 30 minutes always went the same way. Someone would draw six circles on the whiteboard, someone else would say "the thumb can't reach that," and a third person would counter, "move them up and they cover the minimap." All three were right, and no conclusion came. At the next meeting, the same whiteboard got drawn again.

The problem is that drawing layout drafts and checking whether those drafts follow the rules are tangled together inside one person's head. The person who drew a draft has a hard time rejecting it. This chapter pulls those two apart. The tedious work of drafting multiple layouts goes to AI, and code rejects any draft that violates the overlap, thumb-corner, and touch-size rules. The human stands only at the spot where, among the drafts the code has passed, one gets picked for "game feel." If 9.1 built the rulebook for the whole HUD (heads-up display), this chapter is one full cycle of applying that rulebook, end to end, to skill buttons — the single piece your hands touch most.


9.2.1 Why Skill Buttons Are Hard — Information You Press, Not Information You Read

Most elements on a HUD are read-only. Nobody taps the HP bar. That is why, in the thumb-corner diagram of §9.1, HP, MP, and target health could sit in the unreachable top reading zone. Skill buttons are the exact opposite. They must be pressed precisely, on a 0.1-second scale, and in combat your eyes are on the enemy, so your finger finds the position from memory. Shift the position even slightly and a mistap happens right there.

For mobile MMORPGs, the landscape two-handed grip is the standard: pressable elements go in the two bottom corners, and consumables/slots in the bottom center (§9.1 covers why landscape is the standard and what the three zones are). Within that standard, skills almost all land in the bottom-right corner cluster the right thumb can reach (the left thumb is tied up with movement at the bottom left). One distinction applies — active skills pressed on a 0.1-second scale belong in this bottom-right cluster, but consumables, auto-use items, and quick slots go in a separate slot bar at the bottom center, between the two thumbs. This chapter covers active skill buttons only, and every coordinate check assumes the landscape two-handed grip.

So skill button placement is bound by three deterministic rules at once — minimum touch target (44pt per Apple's Human Interface Guidelines, HIG), spacing between adjacent buttons (8dp per Google's Material Design), and thumb reach (skills go in the right thumb's bottom-right corner). All three are already in the rulebook built in §9.1.1 as items judgeable by coordinates and size, so the public-standard numbers follow that rulebook (44pt touch and 8dp spacing are certified figures; only the right-thumb corner is an industry-common model). These three become the primary input to the lint that rejects AI layout drafts in this chapter. When the code says "skill_3 is 40pt, below the HIG 44pt minimum" instead of "isn't this button a bit small?", the 30 minutes at the whiteboard disappear.

Putting the platform baselines side by side with PC makes the starting point clear. PC is precise and high-volume; mobile landscape is restricted to the two-handed corners (see the §9.1 rulebook for the full comparison table). Looking at skill input alone, the difference is plain — on PC, hotkeys let skills sit anywhere on screen: fingers stay on the keyboard, so reach is a non-issue and many slots are possible. Mobile landscape has no hover and no hotkeys, so skills must be laid in the bottom-right corner the right thumb reaches, in frequency order (with a limit of 6\~8 exposed at once), and the most-used skill must sit on the inner side of the corner (the most reachable spot). So the essence of mobile skill layout is not "a pretty arrangement" but "frequency-ordered priority placement inside the right-thumb corner, plus rulebook review." And drawing multiple drafts is tedious when done by hand, and the standard drifts every time. Tedious, fickle repetitive work — exactly the spot where AI outlasts a human.


9.2.2 [Worked Transcript] Layout Drafts for a New Class's Six Skills — Having AI Draft Three Options

I will show one full cycle — from input through draft rejection to the end — of placing the six active skills of the new "Shaman" class on mobile. What follows faithfully reproduces a new-skill UI session from my own project (a mobile-first MMORPG, "Project A" below). The inputs and prompts can be copied as-is; the outputs are reconstructed from the actual session.

Step 1 — Input: Skill Specs as a Machine-Readable Table

Turn the six skills' use frequency and basic character into yaml. The use rates are values pulled from combat logs in the data sheet, so nothing here is newly made up.

# skill_set_shaman.yaml — new class 'Shaman', 6 active skills
screen: { w: 2400, h: 1080, dpr: 3 }   # based on a 6.x-inch landscape screen, pt = px / dpr
skills:
  - id: s1_quickbolt    # basic attack, most frequent
    use_rate: 0.41      # share of use in combat (extracted from logs)
    role: spam          # rapid tapping
  - id: s2_hex          # debuff, frequent
    use_rate: 0.22
    role: core
  - id: s3_totem        # placeable, moderate
    use_rate: 0.14
    role: core
  - id: s4_heal         # heal, occasional but urgent
    use_rate: 0.11
    role: panic         # instant in emergencies
  - id: s5_curse        # AoE debuff, occasional
    use_rate: 0.08
    role: situational
  - id: s6_ultimate     # ultimate, rare
    use_rate: 0.04
    role: burst

The key slots are use_rate and role. The most-pressed s1_quickbolt (41%) and s4_heal (panic) — which must be found within 0.2 seconds in an emergency — have to sit where the right thumb reaches best (the inner bottom-right corner). The rarely used s6_ultimate (4%) can sit at the corner's edge, a bit farther out. This priority is the entire input for the AI layout in the next step.

Step 2 — Prompt: Force Three Drafts and Get Coordinates as Numbers

The attached yaml is a new class's 6 active skills. Make 3 skill button layout drafts for a landscape two-handed grip screen. Lay the skills in the bottom-right corner the right thumb reaches (the left hand is tied up with movement at the bottom left).
The three drafts must follow different philosophies — e.g., a corner fan, a 2-row grid, frequency-weighted. For every button give x,y,w,h in px (screen 2400x1080, dpr 3, pt=px/3).
Put frequently used or role=panic skills on the inner side of the corner where the right thumb reaches well; touch targets at least 132px (44pt), adjacent spacing at least 24px (8dp).
For each draft you must also write one line of design intent and one weakness you see — a draft with no weakness is not allowed. Per draft, only (layout philosophy|coordinate table|design intent|self-weakness), no explanations.

The part of this prompt I worked on harder than the placement instructions is the part that forces the output format and the weakness disclosure. Without numeric coordinates I cannot run the lint, so I pinned down px output, and I made the AI submit "one weakness of this draft" alongside, so it states its own limits first. Ask it freely to "make a pretty UI" and what comes back is the average layout of a generic mobile RPG, in words, with no picture.

Step 3 — Claude's Output (Reconstructed from the Actual Session, Excerpt)

== Draft A: bottom-right corner fan ==
Layout philosophy: 6 buttons fanned around the right thumb's pivot. Frequency order, inner→outer.
| id          | x    | y   | w   | h   | pt   |
|-------------|------|-----|-----|-----|------|
| s1_quickbolt| 2160 | 912 | 168 | 168 | 56pt |
| s2_hex      | 2016 | 852 | 132 | 132 | 44pt |
| s3_totem    | 2052 | 996 | 132 | 132 | 44pt |
| s4_heal     | 1884 | 924 | 132 | 132 | 44pt |
| s5_curse    | 1920 | 780 | 120 | 120 | 40pt |
| s6_ultimate | 1788 | 840 | 120 | 120 | 40pt |
Design intent: the frequently used s1 goes largest, on the easiest-to-reach inner bottom-right corner.
Self-weakness: s5,s6 at 40pt may fall short of HIG. Right-thumb reach is iffy at the fan's outer edge.

== Draft B: 2-row grid (2x3) ==
Layout philosophy: 2-row x 3-column grid at the bottom right. Visual alignment first.
| id          | x    | y   | w   | h   |
|-------------|------|-----|-----|-----|
| s1_quickbolt| 2208 | 912 | 132 | 132 |
| s2_hex      | 2052 | 912 | 132 | 132 |
| s3_totem    | 1896 | 912 | 132 | 132 |
| s4_heal     | 2208 | 756 | 132 | 132 |
| s5_curse    | 2052 | 756 | 132 | 132 |
| s6_ultimate | 1896 | 756 | 132 | 132 |
Design intent: all buttons the same size; the alignment is clean.
Self-weakness: the upper row (s4,s5,s6, y=756) sits at the upper boundary of the right-thumb corner, so reach is iffy.
         The emergency heal s4 is on the upper row, so it may be slow to find in a crisis.

== Draft C: frequency-weighted fan ==
Layout philosophy: s1 as a large central circle, the rest wrapped around the right thumb's pivot in frequency order.
| id          | x    | y   | w   | h   |
|-------------|------|-----|-----|-----|
| s1_quickbolt| 2184 | 900 | 180 | 180 |
| s2_hex      | 2028 | 852 | 144 | 144 |
| s3_totem    | 2052 | 996 | 132 | 132 |
| s4_heal     | 1908 | 912 | 144 | 144 |
| s5_curse    | 1932 | 792 | 120 | 120 |
| s6_ultimate | 1812 | 972 | 120 | 120 |
Design intent: the hand goes to s1 naturally, and the emergency s4 is enlarged and kept on the inner corner.
Self-weakness: being a fan, button spacing is uneven. Proximity collision risk between s2-s5 and s4-s6.

That all three drafts reported a self-identified weakness is the core of this output. A flagged "possible 40pt shortfall," B "emergency heal on the upper row," C "proximity collision risk." The AI pointed first at the weak spots of its own drawing. But this is only self-reporting; the real verdict comes from the code.

Step 4 — lint: The Code Rejects All Three Drafts

Compare the three drafts by eye and the taste fight starts again — "B looks cleaner, though." So I feed all three, as-is, into skill_layout_lint.py from §9.2.3. The results came out like this.

[Draft A] bottom-right corner fan
  [FAIL] B-size  : s5_curse 40pt < 44pt (below HIG)
  [FAIL] B-size  : s6_ultimate 40pt < 44pt (below HIG)
  [WARN] C-corner: s6_ultimate x=1788 — corner's left boundary, right-thumb reach 'medium'
  → passed 4/6, critical violations 2

[Draft B] 2-row grid (2x3)
  [FAIL] C-corner: s4_heal     y=756 (0.70h) < 0.55h not below → above the right-thumb corner
  [FAIL] C-corner: s5_curse    y=756 (0.70h) < 0.55h not below → above the right-thumb corner
  [WARN] role    : s4_heal(panic) y=756 — emergency skill on the upper row
  → passed 4/6, critical violations 2

[Draft C] frequency-weighted fan
  [FAIL] A-overlap: s2_hex ∩ s5_curse gap 18px < 24px (below 8dp)
  [FAIL] A-overlap: s4_heal ∩ s6_ultimate gap 12px < 24px (below 8dp)
  → passed 4/6, critical violations 2

All three drafts failed. What is interesting is that the self-reports and the lint verdicts nearly overlapped. Violations actually occurred where the AI had called out "weaknesses." But the self-report said "risk," while the lint says a number: "s2_hex and s5_curse, 18px apart." Nothing is left to debate in the meeting.

There is one important decision at this step. All three failing does not mean going back to square one. Paste the lint report into the next prompt, as-is, and re-prompt.

Step 5 — Re-Prompt: Feed the lint Report Back As-Is

Pasting the lint results for all three drafts as-is (the report above). Pick only the one draft whose violations are easiest to fix and give me a single revision that brings every FAIL to 0.
All buttons at least 132px (44pt), spacing at least 24px (8dp), and s4_heal(panic) goes in the spot inside the right-thumb corner that reaches second-best after s1.
Give only the coordinate table, plus one line each on which FAIL was resolved and how.
// (intent: inside the corner = x>70% of width(1680), y>55% of height(594))

The AI picked draft A (the corner fan). Its only violations were "two buttons under size," so it required the least surgery. It enlarged the failing s5 and s6 to 132px, and to recover spacing for the larger buttons, it nudged the fan curve slightly toward the inner corner.

== Draft A' (revised) ==
| id          | x    | y   | w   | h   | pt   |
|-------------|------|-----|-----|-----|------|
| s1_quickbolt| 2160 | 900 | 168 | 168 | 56pt |
| s2_hex      | 2010 | 846 | 132 | 132 | 44pt |
| s3_totem    | 2046 | 990 | 132 | 132 | 44pt |
| s4_heal     | 1872 | 918 | 144 | 144 | 48pt |
| s5_curse    | 1908 | 774 | 132 | 132 | 44pt |
| s6_ultimate | 1740 | 858 | 132 | 132 | 44pt |
Resolved: B-size s5 40→44pt / B-size s6 40→44pt /
     C-corner s6 stays inside the corner at x=1740(0.725w)·y=858(0.79h) →
     role: s4_heal enlarged to 144px to strengthen emergency identification.

I fed draft A' back into skill_layout_lint.py.

[Draft A'] bottom-right corner fan (revised)
  [PASS] B-size  : all buttons ≥ 44pt
  [PASS] A-overlap: minimum gap 30px ≥ 24px
  [PASS] C-corner : all control buttons inside the right-thumb corner (x≥1680, y≥594)
  [WARN] C-corner : s6_ultimate x=1740 — corner's left edge, reach 'medium'
  → passed 6/6, critical violations 0, WARN 1

FAIL reached 0. The one remaining WARN (s6_ultimate sits at the corner's left edge, so right-thumb reach is "medium," not "easy") is not auto-killed by the code. It goes up to a human. And this WARN is in fact intended design. s6 is the ultimate, used least at a 4% use rate, so the innermost corner spot should be ceded to the frequently used s1, and the edge is the right place for it. A human ruled "this WARN is intended" and passed it. The cycle — input → 3 drafts → lint → wipeout → re-prompt → pass — closes here.

This one loop is this chapter's bar for Show. Unless you watch to the end — what the AI draws, what the lint rejects, and which WARN a human keeps alive — the sentence "we generated UI drafts with AI" is hollow.


9.2.3 The lint as Code — Overlap, Thumb Corner, and HIG Size

The heart of the cycle above is some 30 lines of code that reject violations of three rules. The three items in the §9.2.1 table become three functions, one for one.

# skill_layout_lint.py — skill button layout verification (skeleton)
# Input: list of button coordinates from the AI [{id, x, y, w, h, role, use_rate}]
# Output: list of A-overlap / B-size / C-corner violations
# Premise: landscape two-handed grip. Skills go in the bottom-right corner the right thumb reaches.

MIN_TAP_PX    = 132    # HIG 44pt * dpr 3 = 132px
MIN_GAP_PX    = 24     # Material 8dp * dpr 3 = 24px
RIGHT_CORNER_X = 0.70  # right of 0.70 of screen width = right-thumb corner
BOTTOM_Y       = 0.55  # below 0.55 of screen height = bottom corner

def in_right_thumb_corner(b, w, h):
    """Is this inside the bottom-right corner the right thumb reaches in landscape grip?
    (left thumb = bottom-left movement, right thumb = bottom-right skills)"""
    rx, ry = b["x"] / w, b["y"] / h
    return rx > RIGHT_CORNER_X and ry > BOTTOM_Y

def lint(buttons, screen_w, screen_h):
    issues = []
    # Rule B: minimum touch target size (HIG 44pt)
    for b in buttons:
        side = min(b["w"], b["h"])
        if side < MIN_TAP_PX:
            issues.append(f"[FAIL] B-size : {b['id']} {side//3}pt "
                          f"< 44pt (below HIG)")
    # Rule A: adjacent button overlap/gap (distance between the two nearest edges)
    for i, a in enumerate(buttons):
        for c in buttons[i+1:]:
            gap = edge_gap(a, c)          # shortest gap between two rects (px)
            if gap < MIN_GAP_PX:
                issues.append(f"[FAIL] A-overlap: {a['id']} ∩ {c['id']} "
                              f"gap {gap}px < {MIN_GAP_PX}px (below 8dp)")
    # Rule C: control elements stay inside the right-thumb corner. For panic, the more inner the better.
    for b in buttons:
        rx, ry = b["x"] / screen_w, b["y"] / screen_h
        if not in_right_thumb_corner(b, screen_w, screen_h):
            issues.append(f"[FAIL] C-corner: {b['id']} "
                          f"x={b['x']}({rx:.2f}w) y={b['y']}({ry:.2f}h) "
                          f"→ outside the right-thumb corner")
        elif b.get("role") == "panic" and rx < 0.78:
            issues.append(f"[WARN] role   : {b['id']}(panic) "
                          f"emergency skill near the corner's inner boundary")
    return issues

This code is what neutralizes the taste remark "but draft B is prettier" in a meeting. Prettiness is something to argue about after the lint has passed a draft. A draft that draws a [FAIL] from the lint does not enter the build, pretty or not. This applies the HUD lint gate built in §9.1.1 all the way through to skill buttons, the trickiest single piece — and the same division of labor holds here: what can be judged by coordinates and size goes to code, and "is this WARN intended?" goes to a human.

Here is the whole cycle at a glance.

flowchart LR
    A["Skill spec yaml
(use_rate·role)"] --> B["AI: 3 layout drafts
coordinates + self-reported weakness"] B --> C{"skill_layout_lint.py
overlap·size·right-thumb corner"} C -->|has FAIL| D["Re-prompt with the
lint report as-is"] D --> B C -->|FAIL 0, WARN only| E["Human: judge whether
the WARN is intended"] E --> F["Layout finalized
+ ArtGuide 06_UI sync"] classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; class A data; class B ai; class C code; class E human; class F pass;

Human hands touch only two places: the very front, where the input spec goes in clean, and the very end, where the WARNs the lint cannot kill get judged. The tedious three-draft generation and coordinate checking in between are run by the AI and the lint.


9.2.4 Log the Pass Rate — See the Tool's Performance in Numbers

Pull one layout and stop, and you cannot tell whether this tool works. So I log the lint results every time. What gets recorded is simple — how many of the AI's drafts passed the first lint run (first-pass rate), and how many re-prompts it took to reach FAIL 0 (round-trip count).

The figures below are measured values I counted myself while building the skill UIs of three new classes (the Shaman plus two others) through this cycle. The sample is small — 3 classes, 9 layout sessions — so the right way to read them is as directional values, not precise population parameters. None of the numbers are doctored.

Item Measured Note
AI first drafts passing the first lint run 1 of 9 the other 8 had one or more FAILs
Average FAILs on the first run 1.8 per draft mostly size shortfalls or outside the right-thumb corner
Average round trips to FAIL 0 1.4 lint-report re-feed method
Most common FAIL type B-size (size shortfall) C-corner (right-thumb corner) next

The most important row is the first one. The AI's first drafts failed the lint 8 times out of 9. That is not the tool failing — it is the signal of normal operation. Let the AI emit coordinates freely and it violates HIG 44pt often. The lint catches it every time, and feeding the report back reaches 0 in one or two round trips. If the first-pass rate had been 100%, that would mean the lint is too loose — not that the AI is perfect.

This pass-rate log also becomes the basis for deciding whether to tighten or loosen lint rules. If some FAIL type gets released by human hands every time as "actually intended," that rule is too strict. Conversely, if mistap complaints come in after launch on layouts the lint passed, the rules are too loose.


9.2.5 The Final Layout as a Picture — Button Layout SVG

Drawing the lint-passing draft A' from §9.2.2 at its exact coordinates gives the picture below. What the table's numbers look like on an actual screen only sinks in as a picture. Holding a landscape phone with both hands, the left thumb lands at the bottom left (movement) and the right thumb at the bottom right (the skill cluster). Circle size is proportional to the touch target (pt), and color is thumb-reach difficulty (green easy / yellow medium).

Top — status display only (HP · MP · target, read-only) Game view (where combat happens) Move L thumb Right-thumb 'easy' corner ↘ HP MP Tgt Map s1 56pt s2 44 s3 44 s4 panic s5 44 s6 medium easy medium (s6 = rare ultimate, intended)

Seen as a picture, the lint report's final WARN makes sense at a glance. Only s6_ultimate (yellow) sits at the left edge of the bottom-right corner, a "medium" right-thumb-reach spot. But s6 is the ultimate with a 4% use rate, so the corner's edge is the right place for it. The most-used s1 (green, 56pt, the largest) sits in the inner bottom-right where the right thumb reaches best, and the emergency heal s4 (yellow border) is enlarged so the hand finds it fast in a crisis. The left thumb is tied to "movement" at the bottom left, so the skills all gather in the right corner. One coordinate table matching one picture exactly — that is why I demanded coordinates as numbers.


9.2.6 Common Failures

Pattern Why it fails Fix
Drawing only circles on a whiteboard, then meeting No coordinates, so no lint — the taste fight repeats Get coordinates in px and feed them to the lint (§9.2.2)
Wholesale delegation: "AI, make a pretty skill UI" Without a rulebook, you get the average generic RPG layout A prompt that forces 3 drafts + coordinates + self-reported weaknesses
Laying out for a portrait one-handed grip MMORPGs standardize on landscape two-handed; skills go in the right-thumb corner Lint against landscape 2400x1080 and the bottom-right corner
Comparing drafts by eye only HIG shortfalls and overlaps get missed every time Automatic verdicts with skill_layout_lint.py
First draft passes the lint → relax, the tool works Can be a sign the lint is loose Check rule tightness with the pass-rate log (§9.2.4)
Code auto-blocks WARNs too Kills intended placements (a rare ultimate) as well WARNs go to human judgment (§9.2.3)

The fifth one gets missed most often. It feels good when the AI's first draft passes every time, but that usually means the lint rules are slack. Failing 8 times out of 9 is the healthy state.


9.2.7 Try It Yourself — One Step You Can Take Today

Solo Scale-Down: You don't need the lint code. Pick 4\~6 skills from your own game (or a game you love), write the spec by hand in the §9.2.1 format (for use_rate, a rough frequency ranking is enough), paste the §9.2.2 prompt as-is, and get three drafts back. Then, instead of a ruler, keep just "44pt = 132px" in your head and circle, by hand, every button under 132px in the AI's coordinate table. Then, treating it as a landscape screen, check whether any skill falls outside the bottom-right corner (right of 70% of the width, below 55% of the height). That one pass teaches you, hands-on, what the lint does.

If you're on a team, start with this one step. Pin down the three functions of §9.2.3's skill_layout_lint.py (size, spacing, right-thumb corner) in code first. Three functions are enough. With the rulebook in place, AI drafts and designer mockups get measured against the same line, and only lint-passing drafts move on to the art team's 96_ArtGuide/06_UI/, auto-synced along the _convert_md_to_html.py_SyncToArtRepo.bat path. Until the finalized coordinates reach the art team, the human's last job is a single ruling on one WARN: "this is intended."


Key Takeaways

Next Chapter Preview

9.3 ArtGuide/06_UI Collaboration — Designers Write md; the Art Team Sees Only html

Primary reader: the UX/UI designer who collaborates daily with a non-design discipline (art) on a mid-size team Scaled-down version for solo/hobbyist readers: §9.3.8, "If You're Solo, Just This Much"

When a designer keeps UI decisions in Markdown, work gets clean. You get version control, visible diffs, and a document you can throw at an AI as-is. The problem is that the art team doesn't read Markdown. More precisely, they have no reason to. Tell an artist "grab 아트_결정사항.md (the art-decisions file) from SVN and take a look," and half of them haven't installed an SVN client, while the other half open it in Notepad, stare at broken ## headers and table syntax, and ask, "How am I supposed to read this?"

The wrong prescription here is "teach the art team Markdown." An artist's time should go into pushing pixels. Every hour spent learning Markdown conventions, SVN checkouts, and how to read a diff is pure loss. The right prescription is to automate conversion and delivery on the design side, so that the art team's learning burden drops to zero. The designer writes md, a script converts it to html, another script pushes it into the art repository, and the art team sees only html in a browser. This chapter runs that pipeline end to end, once — from the seat where AI drafts the decisions, through the automation of conversion and delivery, to what the human actually rejects.


9.3.1 Where Collaboration Really Breaks Is the 'Format'

Plenty of books explain design–art friction as "ambiguous decision rights" — who decides the color, who decides the function. That split matters, but no matter how well you draw the responsibility chart, if the art team can't read the chart, nothing happens. In practice, the place where accidents actually occur is not decision rights but the delivery format.

These are the accidents that actually kept recurring on my project (a mobile-first MMORPG, "Project A" hereafter).

Accident Surface cause Real cause
Art worked from an outdated decision doc "I never got the latest one" Delivery was manual (email attachment), so it slipped through
The decision table looked broken "Why does it look like this?" The md was opened in Notepad
"Where is that decision written down?" Verbal handoff The canonical record was scattered across chat

None of the three is a decision-rights problem. They all happen because the canonical document is not delivered in a format the art team reads, automatically, always up to date. That is why this chapter's tool is a delivery pipeline, not a responsibility chart. The split of responsibilities is agreed once and done; delivery has to happen every single time a decision changes.

Start with the actual folder structure. Project A's art guide lives under workspace/96_ArtGuide/, split into 7 domains.

96_ArtGuide/
├── 00_Common/      # Shared standards (style, color palette, lighting)
├── 01_Character/
├── 02_Animation/
├── 03_Monster/
├── 04_NPC/
├── 05_VFX/
├── 06_UI/          # ← the domain this chapter covers
└── 07_Env/

(In the comments: 00_Common holds shared standards — style, color palette, lighting; 06_UI is the domain this chapter covers.)

And two operational files live in this folder alongside the content: _convert_md_to_html.py and _SyncToArtRepo.bat. These two files are the spine of this chapter.


9.3.2 The Four-Stage Sync Pipeline — From the Designer's md to the Art Team's Browser

The whole flow has four stages. The key point: the human (the designer) touches only the md in stage 1; the remaining three stages are run entirely by scripts. The art team sees only the html in stage 4. They don't even need to know the md exists.

flowchart LR
    A["Stage 1 · Design team SVN
아트_결정사항.md
(written by the designer / AI draft)"] A --> B["Stage 2 · _convert_md_to_html.py
md → html conversion
(tables, headers, image embeds)"] B --> C["Stage 3 · _SyncToArtRepo.bat
auto-push to the art SVN
(separate repository)"] C --> D["Stage 4 · Art team browser
views html only
(zero md-convention learning burden)"] D -.feedback / change requests.-> A classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef human fill:#fde68a,stroke:#b45309,color:#000; class A,D human; class B,C code;

Here is exactly what each stage does.

Stage 1 (designer, human) — Write the decisions in Markdown in 06_UI/아트_결정사항.md. How to put AI into this seat is the spine of §9.3.4. Decisions are items like "button primary color #3A7BD5" or "minimum touch target 44pt."

Stage 2 (_convert_md_to_html.py, automated) — Converts md to html. Not a bare conversion: it renders tables so the art team can actually read them, embeds ![](...) image references inline, and attaches a table of contents. Out comes a self-contained html file an artist can open in a browser with a single double-click.

Stage 3 (_SyncToArtRepo.bat, automated) — Pushes the converted html to the art team's separate SVN repository. The separation of the design repository and the art repository is the point. The art team only ever looks at its own repository, and never needs to know the permissions or structure of the design one.

Stage 4 (art team, human) — The artist opens the synced html from their own repository in a browser. No Markdown syntax, no SVN commands, no diff-reading to learn. Zero md-convention learning burden is both the design goal and the success criterion of this pipeline.

Feedback flows from stage 4 back to stage 1. When art says "this decision looks off," the designer fixes the md, and stages 2–3 run again automatically. Art only has to reopen the updated html.


9.3.3 Why Automate Conversion and Delivery — The Asymmetry of Learning Burden

Pause here and make the design intent explicit. Converting md to html is trivial in itself. The real design lies in deciding who absorbs whose learning burden.

There were two options.

Option A — Art team learns md Designer writes md only 5 artists ×SVN·md learning Learning cost = one authoring pass × multiplied by artist headcount The burden eats into pixel time → rejected Option B — Designer automates Designer md + script, once 5 artists double-click html Learning cost = designer, once (zero burden on art) Art stays focused on pixels → persists

(In the diagram: option A — the art team learns md; the learning cost is one authoring pass multiplied by the number of artists, the burden eats into pixel time, and the scheme gets rejected. Option B — the designer automates; the learning cost is one designer, once, with zero burden on art; artists stay focused on pixels, and the scheme persists.)

The point is the asymmetry. In option A, the learning cost multiplies by the number of artists, and it recurs with every new hire. In option B, the designer writes the script once and the marginal cost on the art side is zero. Pile the burden onto the side that can automate, not the side with more people — that is rule number one for collaboration tools that cross into a non-design discipline. When this rule breaks — when a collaboration tool forces new learning on the other discipline — that tool stops being used within a quarter or two.


9.3.4 [Worked Transcript] Drafting the UI Decisions md with AI

I said the designer writes the md in stage 1; here is one full cycle of using AI to draft that md. After a decision meeting you're left with scattered notes — chat messages, whiteboard photos, verbal agreements. Turning them into a canonical decisions md is tedious, and the format drifts every time. That is exactly the kind of job AI is for. The boundary that matters: the human makes the decisions; the AI only arranges those decisions into a fixed format.

Step 1 — Input: The Raw Meeting Memo

[UI decision meeting memo — 06_UI skill slots, raw]
- Agreed to enlarge the skill slot buttons. Too small on mobile.
- Art picks the colors. But primary tone stays in the blue family.
- Slot disabled (cooldown) state agreed as gray + number overlay
- Localization... what to do about long skill names? On hold for now
- Oh and long-press should pop the skill description (this is a feature, design confirmed)

In translation, the memo says: we agreed to enlarge the skill-slot buttons — they're too small on mobile; art will pick the colors, but the primary tone stays in the blue family; the slot's disabled (cooldown) state was agreed as gray plus a number overlay; what to do about long localized skill names — on hold for now; and long-press should pop the skill description (a feature, confirmed by design).

Step 2 — The Prompt: Forcing the Decision/Hold/Owner Split

Attached is the raw memo from the UI decision meeting. Organize it into a decisions Markdown to hand to the art team.
Classify each item as [확정]/[보류]/[아트결정] (확정 = confirmed: design decided, art follows; 보류 = on hold: undecided + reason; 아트결정 = art's call: visual treatment, art decides and design supplies constraints only), and attach one line of rationale/constraint per item — if the memo has no rationale, honestly write '근거 미기재' (no rationale recorded) and don't fill it in with guesses.
Where something can be verified against a public standard, like touch sizes, name the standard, and never invent items that are not in the memo.
Use the table columns |Item|Class|Decision/Constraint|Rationale|Verifiable|, include frontmatter (title, layer: L1), and give me the full document.

In translation, the prompt says: "Attached is the raw memo from the UI decision meeting. Organize it into a decisions Markdown to hand to the art team. Classify every item as [확정]/[보류]/[아트결정] — confirmed (design decided, art follows), on hold (undecided, with the reason), or art's call (visual treatment art decides; design supplies constraints only) — and attach one line of rationale/constraint per item. If the memo has no rationale, write 'no rationale recorded' honestly instead of filling the gap with guesses. Where a public standard applies, like touch sizes, name the standard, and never invent items that are not in the memo. Use the table columns |item|class|decision/constraint|rationale|verifiable|, include frontmatter (title, layer: L1), and give me the full document."

The weight of this prompt sits not on the cleanup work but on forced classification and a ban on invention. Ask an AI to organize decisions and it will helpfully add "best practices" that were never in the memo. To the art team, those additions read as "decisions design has confirmed." Hence the nails this prompt drives in: the [확정]/[보류]/[아트결정] classification, the rationale requirement, and "if it's not in the memo, don't make it up."

Step 3 — Claude's Raw Output (Reconstructed from the Actual Session)

---
title: 06_UI Decisions — Skill Slots
layer: L1
---

# 06_UI Decisions — Skill Slots

| Item | Class | Decision/Constraint | Rationale | Verifiable |
|---|---|---|---|---|
| Skill slot button size | [확정] | minimum 44×44 pt | mobile touch target, Apple HIG 44pt recommendation | ✅ coordinates/size |
| Primary tone | [아트결정] | keep the blue family (constraint), exact value is art's | memo: "primary tone stays in the blue family" | △ color value |
| Cooldown state treatment | [확정] | gray overlay + remaining number | meeting agreement | ◯ visual review |
| Multilingual skill-name length | [보류] | later meeting | memo: "On hold for now" — long languages (German/Russian) untested | — |
| Long-press skill description | [확정] | long-press → tooltip | confirmed feature (design territory) | ◯ behavior review |
| Slot spacing | [확정] | minimum 8 dp recommended | 근거 미기재 (no rationale recorded) — not in the memo; Material recommended value for reference | ✅ |

In translation, the table rows read: skill-slot button size — [확정] minimum 44×44 pt (mobile touch target, Apple HIG 44pt recommendation; verifiable by coordinates/size). Primary tone — [아트결정] keep the blue family (constraint), exact value is art's (color value partially verifiable). Cooldown state — [확정] gray overlay plus remaining number (meeting agreement; visual review). Multilingual skill-name length — [보류] for a later meeting (memo: "on hold for now" — long languages like German/Russian untested). Long-press skill description — [확정] long-press shows a tooltip (confirmed feature, design territory; behavior review). Slot spacing — [확정] minimum 8 dp recommended; rationale: "no rationale recorded — not in the memo; Material recommended value for reference."

Step 4 — Verification and Rejection (The Human's Job)

This output must not go up as the canonical version as-is. One line snags.

It's the last row, 슬롯 간 간격 | 최소 8 dp (slot spacing | minimum 8 dp). Go back to the memo: spacing never came up in the meeting. The AI judged "a mobile UI ought to have a spacing standard too" and helpfully added the Material Design recommended value. That violates rule 5 ("never invent items not in the memo"). The AI did honestly write 근거 미기재 (no rationale recorded), but it should never have created the row at all. If this one line reaches the art team, it reads as "design confirmed 8dp spacing."

So I push back.

'Slot spacing' is not in the meeting memo — you added it. Remove it from the table.
Items that look like they need a decision but aren't in the memo go only as candidates under a '## 미결 — 다음 회의 안건' (open items — agenda for the next meeting) section at the bottom, not in the table; the decisions table keeps only items that were actually in the memo.

In translation: "'Slot spacing' is not in the meeting memo — you added it. Remove it from the table. Items that look like they need a decision but aren't in the memo go only as candidates under an '## 미결 — 다음 회의 안건' (open items — agenda for the next meeting) section at the bottom; the decisions table keeps only items that were actually in the memo."

The AI removed the spacing row from the table and split the candidates out at the bottom: "Agenda for the next meeting: slot spacing standard (currently undecided), multilingual skill-name length handling." Now the decisions table holds only what the meeting actually decided, and the AI's reasonable candidates have been demoted from "confirmed" to "agenda." This separation matters because when the document the art team receives mixes what is confirmed with what is still under discussion, art will take an undecided item as confirmed and start working on it.

That one round trip completed stage 1 (the md). Now it leaves human hands and moves to the stage 2–3 automation.


9.3.5 Stages 2–3 Automation — Humans Don't Touch Conversion or Delivery

The finished md is now handled by scripts. The skeleton of the conversion script is simple.

# _convert_md_to_html.py (skeleton)
# Input: 06_UI/*.md (decisions written by the designer)
# Output: .html with the same name (a self-contained file the art team opens in a browser)

def convert(md_path):
    md_text = read(md_path)
    front, body = split_frontmatter(md_text)          # extract title and layer
    html_body = markdown_to_html(body, extensions=[
        "tables",        # table rendering (fixes the broken tables art saw in Notepad)
        "fenced_code",
    ])
    html_body = embed_images_inline(html_body, base_dir=md_path.parent)
    # ↑ inline-embeds references like ![](char_skill_ui.png) →
    #   art doesn't have to fetch image files separately
    toc = build_toc(html_body)                         # auto-generate the table of contents
    return render_template(title=front["title"], toc=toc, body=html_body)

The point is that the conversion isn't a bare md→html pass. It does three more things. It renders tables properly (the broken |---| the artist used to see in Notepad disappears), it embeds images inline (the artist doesn't have to fetch image files separately), and it auto-generates a table of contents (when the decision doc grows long, the artist can jump straight to the item they want). These three are what actually make "just look at the html" hold.

The delivery script ties it together like this.

REM _SyncToArtRepo.bat (skeleton)
REM 1) Convert all md in 06_UI to html
python _convert_md_to_html.py 06_UI\*.md

REM 2) Copy the converted html into the art SVN working copy
xcopy 06_UI\*.html %ART_REPO%\UI\ /Y

REM 3) Auto-commit and push to the art SVN (separate repository)
svn add %ART_REPO%\UI\*.html --force
svn commit %ART_REPO%\UI -m "[auto] 06_UI decisions updated"

All the designer does is double-click _SyncToArtRepo.bat once (or wire a hook so it runs automatically when a decisions commit lands). Conversion, copying, and the push to the art repository run in one go. The art team updates their repository and the latest html is waiting.

How far does AI go here — You can have AI write this stage 2–3 automation code. "Write a script that takes a folder of md, converts it to html with tables and images included, and pushes it to a separate SVN" is squarely in AI's wheelhouse. But which decisions to mark as confirmed, and what to hand over as art's call (§9.3.4), is never delegated to AI. Code goes to the AI, decisions stay with the human — the division this book repeats throughout applies unchanged here.


9.3.6 Image Prompts Also State 'Design Intent' First

The classic way AI gets misused in art collaboration is image prompts. Designers reach for image-generation AI when handing references to the art team or when visualizing a concept quickly. The common mistake is starting from a description of the result ("blue rounded button, glow effect, 4K").

One of my collaboration principles is image_prompt_design_intent_firsteven an image prompt states the design intent first, not a description of the result.

Approach Prompt Problem / effect
Result-first (bad) "Blue rounded button, glow, 4K, game UI" Art can't ask "why blue?" — the intent evaporates
Intent-first (good) "A skill button that makes the cooldown state instantly legible. Active = a visual pull that makes you want to press it right now; on cooldown = restraint. Tone stays in the primary blue family" Art reads the intent and can counter-propose a better visual

The difference is what the art team can do when they receive the prompt. Given only a result description, art either draws it as written or ignores it. Given the design intent, art can propose its own visual that solves that intent better. This is how a designer gives direction without trespassing on art's decision territory (the [아트결정] class from §9.3.4). The designer supplies the "what for"; art decides the "how it looks."

So when an image reference goes into the §9.3.4 decisions md, the caption reads not "blue button" but "a slot whose purpose is distinguishing the cooldown state — exact treatment is art's call." The conversion script embeds the caption together with the image into the html, so art receives the image and the intent as one.


9.3.7 Measurement — What Can Be Counted Honestly

There is a temptation to write this pipeline's effect as a number like "collaboration accidents dropped 70%." Numbers like that, unverified, eat away at the book's credibility. So I separate things honestly.

What public standards can verify — The public-standard values that ride along in the decisions — 44pt touch targets, 8dp spacing, 4.5:1 contrast — follow the §9.1 rulebook. They are not invented numbers; they are cited as-is and can be auto-checked with lint.

Operational metrics that can be measured — What this pipeline can actually count is this: the number of incidents where art worked from an outdated version (with automated delivery this converges to 0), the time it takes a new art-team hire to open the decisions for the first time (a double-click on html puts it on the order of minutes), and the lag between a decision change and its arrival in the art repository (the script's run time). These three are countable from logs and observation, not from "feel."

Author's estimate (unverified hypothesis) — "Fewer misses than in the manual-email days" is clearly the direction, but I did not keep a separate sample, so I won't claim an exact reduction rate. Read it as a direction rather than an absolute value: when delivery hangs on human hands, something will slip in a busy week without fail; when delivery is a script, the misses disappear structurally.


9.3.8 Try It Yourself — One Step You Can Take Today

If you're solo, just this much: You don't need an art team or SVN. Imagine you're handing UI decisions to a contract artist you commission, or a friend you collaborate with. Take the §9.3.4 prompt as-is and have AI produce a one-page md that sorts the scattered UI decisions in your head into [확정]/[보류]/[아트결정] (confirmed / on hold / art's call). Then find one item the AI "helpfully added" (something not in your notes) and push back: "I never decided this — take it out." That is when the boundary between human and AI in decision cleanup sinks in physically. For conversion, the markdown package is enough: one line, python -m markdown decision.md > decision.html.

If you're on a team, start with this one step. Don't begin by building grand two-way sync; put in one line of conversion plus one line of delivery first. A conversion script that turns the decisions md into html (just the table rendering and image embedding from §9.3.5), and one line that copies the html to wherever art looks — a shared drive or a separate repository. These two lines alone eliminate the most common accident: art opening the md in Notepad and hitting a broken table. The responsibility chart and decision rights come after that.

Summarized as setup → prompt → verify:

Step What to do
setup Put in _convert_md_to_html.py (conversion) plus one line of delivery (copy/push) first
prompt Use the §9.3.4 prompt to organize meeting memos into an [확정]/[보류]/[아트결정] md
verify Reject the items the AI invented (not in the memo) → conversion and delivery run automatically → art checks only the html

Key Takeaways

Next Chapter Preview

Part 10 · Qa Design

10.1 The Integrity-Check Atom — A Cascade That Guards FKs Across 30 Sheets

Friday, 6:40 p.m. We were due to add 12 new quests to the internal build the following Monday. I appended new rows to quest_table, filled in the matching rows in the reward sheet, and wired NPC lines into the dialogue sheet. Three sheets, about 50 rows. I scanned them twice with my own eyes, and everything looked fine.

Monday morning, the build broke. One of the new quests referenced a reward_id that did not exist in the reward sheet. On Friday evening I had deleted a reward row and re-added it, mistyping a single character in the id. rwd_q318 became rwd_q381. It is exactly the kind of typo human eyes can never catch. The two sheets live in different folders and are touched by different people at different times. At 50 rows, eyes can still catch it. Once more than 30 sheets start referencing each other through foreign keys (FKs), the human eye is no longer a verification tool.

This chapter follows one session I actually ran, showing how a single kind of check atom — integrity_check_fk — verifies FK integrity across more than 30 sheets and, when something breaks, notifies the owner through our collaboration tool (a SaaS for managing tasks and schedules — this project uses ClickUp; JIRA and Redmine fill the same seat).

Finding the one line that is off in something someone else built is how I first entered this industry. My first job was QA and review on single-player games, and back then my hands and eyes were the only verification tools I had. Twenty-odd years later, I hand the same job to code — in the very place where the human eye stopped being a verification tool.


10.1.1 What the Check Must Catch — The Anatomy of a Broken FK

First, a picture of what is being checked. Game data sheets work like a relational database. A column in one sheet points to another sheet's primary key. When that arrow breaks, the game dies at runtime — or worse, quietly displays an empty value.

quest_table quest_id (PK) reward_id (FK) npc_id (FK) reward_table reward_id (PK) npc_table npc_id (PK) npc_id 'npc_307' → not in npc_table (broken FK)

The solid green line is a live reference. The value that quest_table.reward_id points to actually exists in reward_table.reward_id. The dashed red line is a dead reference — some quest's npc_id points to an id that is not in npc_table. That red dashed line is what integrity_check_fk catches.

Project A's data sheets have far more than one or two references like this. More than 30 sheets point at each other, and if you draw every FK arrow, the diagram no longer fits on a single page. Tracing them by hand, every time, is impossible.


10.1.2 One Session — How the Check Catches a Broken FK, End to End

From here on, this is a session I actually ran. I am not summarizing: the full prompt, the raw output, and the parts I rejected and re-requested are recorded as they happened.

Step 1 — Have the Checker Read the Sheets and Extract FK Candidates

First I had the checker build the FK map — which column points to which sheet. Registering the FKs of 30 sheets one by one, by hand, is unrealistic, so I had it infer candidates automatically from the column naming convention in the spec sheet (*_id points to the sheet of the same name). The full prompt I entered was the following.

Read the attached spec sheet (the $스키마 tab) and infer which columns in every data sheet look like foreign keys (FKs). Rules: (1) the column name has the form xxx_id, and (2) a sheet named xxx_table or xxx_list has a PK (primary key) column of the same name — then treat it as an FK candidate. Output the inferences as a JSON list of {source_sheet, source_column, target_sheet, target_column, confidence}. For anything the rules cannot settle, mark confidence as low and explain why.

Part of Claude's raw output (copied verbatim, ambiguous items included):

[
  {"source_sheet":"quest_table","source_column":"reward_id",
   "target_sheet":"reward_table","target_column":"reward_id","confidence":"high"},
  {"source_sheet":"quest_table","source_column":"npc_id",
   "target_sheet":"npc_table","target_column":"npc_id","confidence":"high"},
  {"source_sheet":"quest_table","source_column":"prev_quest_id",
   "target_sheet":"quest_table","target_column":"quest_id","confidence":"medium",
   "note":"Name is prev_quest_id, not quest_id, so the rule does not catch it. Presumed self-reference within the same sheet. Needs human confirmation"},
  {"source_sheet":"reward_table","source_column":"item_id",
   "target_sheet":"item_table","target_column":"item_id","confidence":"high"},
  {"source_sheet":"dialogue_table","source_column":"speaker_id",
   "target_sheet":"npc_table","target_column":"npc_id","confidence":"low",
   "note":"Unclear whether speaker_id points to npc_id or to player. Some rows appear to have speaker as 'player'. Cannot be settled"}
]

Two of these items did not get waved through. I adopted prev_quest_id (a self-reference) as a check target, but added an exception by hand: an empty value (the first quest in a chain) is not a violation. I rejected dialogue_table.speaker_id — rows where speaker is 'player' actually exist, so checking this column as an FK would flag every player line as a false violation.

That rejection matters. Use an auto-inferred FK map without human review, and the first run dumps hundreds of false violations on you. The check earns trust only after a human has filtered the map once.

Step 2 — Define the Check Atom from the Reviewed FK Map

I pinned the filtered FK map as the input to the integrity_check_fk atom. The atom format is below. This is the full text of one check atom actually used on Project A.

---
name: integrity_check_fk
description: Verify that every source column value exists in the target sheet's PK, per the registered FK map
type: integrity_check
category: data
priority: P0          # broken FKs block the build
execution_time:
  - on_save           # on sheet save, that sheet only
  - on_build          # all FKs at build time
  - nightly           # full run + report at midnight daily
input:
  fk_map: fk_map.reviewed.json   # the map human-reviewed in Steps 1-2
output_format: violation_list
on_violation:
  - notify: clickup           # notify ClickUp on failure
related_atoms:
  - integrity_check_clickup_notify
  - integrity_check_id_uniqueness
---

The check logic itself is not long. It is a set-membership test: confirm that each value in the source sheet exists in the target sheet's PK set.

def check_fk(fk_map, sheets):
    violations = []
    for fk in fk_map:
        pk_set = {r[fk["target_column"]] for r in sheets[fk["target_sheet"]]}
        for i, row in enumerate(sheets[fk["source_sheet"]]):
            val = row[fk["source_column"]]
            if val in ("", None):          # empty FK is an exception (rule set in Step 1)
                continue
            if val not in pk_set:
                violations.append({
                    "fk": f'{fk["source_sheet"]}.{fk["source_column"]}',
                    "row": i + 2,          # 1 header row + 1-index
                    "value": val,
                    "target": fk["target_sheet"],
                    "severity": fk.get("severity", "P0"),
                })
    return violations

Step 3 — Run the Check and Catch the FKs That Are Really Broken

I ran the check across all 30 sheets with the reviewed map. The output is the standard violation_list. What follows is the actual result from that day (ids and sheet names anonymized; the violation counts and structure are real).

{
  "check": "integrity_check_fk",
  "executed_at": "2026-05-18 09:14:02",
  "input_files": 31,
  "violations": [
    {"fk": "quest_table.reward_id", "row": 318, "value": "rwd_q381",
     "target": "reward_table", "severity": "P0",
     "message": "reward_id 'rwd_q381' not in reward_table. Presumed typo of 'rwd_q318'"},
    {"fk": "quest_table.prev_quest_id", "row": 502, "value": "q_0500",
     "target": "quest_table", "severity": "P0",
     "message": "prev_quest_id 'q_0500' not in quest_table. Presumed notation mismatch (zero-padding) with 'q_500'"}
  ],
  "summary": {"fk_checked": 23, "rows_scanned": 4117, "violations": 2, "passed": 4115}
}

The Friday-evening typo (rwd_q381) was caught on the first line. The second was a different problem I had not known about. One quest's prev_quest_id was q_0500, while the actual quest id was q_500. A notation mismatch caused by zero-padding. To the human eye the two look identical, but as strings they are different values, and the game, unable to find the prerequisite quest, leaves that quest locked. Had it shipped, it is the kind of defect that generates player support tickets.

The "likely a typo" and "likely zero-padding" notes in the message field are where I had the checker go beyond a plain membership failure and also suggest the nearest PK value (by edit distance). It cuts the time a human spends tracing "why did this break?" These guesses are strictly hints, though — the actual corrected value is decided by a human.


10.1.3 When It Breaks — The Cascade up to Collaboration-Tool Notification

That is the behavior of one check. But a check that catches a violation nobody looks at means nothing. The point is the flow that puts the violation directly in front of its owner. On Project A, a separate atom named integrity_check_clickup_notify owns that flow (with an impact score of 294.93 in the JIT metadata, it is one of the highest-rated atoms in the verification bundle — which says that getting an integrity failure to a human matters as much as the check itself).

The full cascade looks like this. The check atoms run in sequence, and a P0 violation at any stage flows into the notification atom.

flowchart TD
    A[Sheet save / build trigger] --> B[integrity_check_id_uniqueness
PK duplicate check] B -->|duplicates found P0| F[Build blocked] B -->|pass| C[integrity_check_fk
FK-map-based integrity check] C -->|0 broken FKs| D[integrity_check_range
reward and numeric range check] C -->|broken FKs P0| E[integrity_check_clickup_notify] D -->|range violations P1| E D -->|pass| G[Checks PASS · build proceeds] E --> H{severity?} H -->|P0| I[Create collaboration tool task
+ mention owner + block build] H -->|P1| J[Collaboration tool comment + alert
build proceeds] I --> F classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; classDef fail fill:#fee2e2,stroke:#dc2626,color:#7f1d1d; class A data; class B,C,D,E,H,I,J code; class G pass; class F fail;

Two design decisions are baked into this cascade.

First, the PK duplicate check runs before the FK check. The FK check assumes that the target sheet's PKs are unique. If the PKs contain duplicates, the question "is this value in the PK set?" stops meaning anything. So the integrity_check_fk atom declares related_atoms: integrity_check_id_uniqueness, and the cascade fixes the order. If the dependency check fails, the FK check is skipped — running it would produce nothing but false results.

Second, the strength of the notification splits on severity. P0 (a broken FK) creates a task in the collaboration tool, mentions the owner registered in the FK map (for reward_table, the rewards designer), and blocks the build. P1 (a reward value outside its recommended range — something that needs review rather than something wrong) leaves only a comment and an alert, and the build goes through. Make every violation a build blocker, and people soon learn to ignore build blockers. Blocking is reserved for what truly must be blocked.

The body of the task actually created in the collaboration tool is one violation_list entry, converted as-is.

[P0] integrity_check_fk violation — build blocked
Sheet: quest_table  |  Column: reward_id  |  Row: 318
Value 'rwd_q381' is not in reward_table.
Nearest candidate: 'rwd_q318' (edit distance 1)
Owner: @rewards_owner  |  Detected: 2026-05-18 09:14  |  Build: nightly-0042

Not a single hand touches the path from check result to a person's inbox. Check → classify → create task → mention is one pipeline. What makes this possible is that violation_list is a standard output format. Whichever check atom caught the violation, the output structure is the same, so one notification atom receives and handles the results of every check.


10.1.4 Operating to Reduce False Violations — Leave Review Evidence

Turn a check on for the first time and false violations are guaranteed. The speaker_id from Step 1 is one example. Leave them unaddressed and people learn to treat the violation report as "mostly false anyway, safe to skip" — the most common path by which a checker loses its credibility.

Project A blocks that path with a principle named human_review_attestation_evidence_mandatory. When you rule something a false violation and carve out an exception, you must leave evidence of who made that call, when, and why. Every exception entry in the FK map file (fk_map.reviewed.json) carries the following.

{
  "source_sheet": "dialogue_table", "source_column": "speaker_id",
  "excluded": true,
  "review": {
    "by": "Lee Minsoo", "at": "2026-05-18",
    "reason": "speaker_id holds either an npc_id or the 'player' literal. Unsuitable for a single FK check.",
    "follow_up": "Consider reintroducing as a branched check after adding a speaker_type column"
  }
}

Without this evidence, when the question "why isn't this column checked?" resurfaces much later, there is nothing to answer it with. So the column goes back into the check, and the hundreds of false violations come back with it. Review evidence keeps the same argument from being fought twice.


Try It Yourself — Setting Up Your First FK Integrity Check

A minimal procedure for readers who want to bring FK checking to their own data sheets.

setup. Gather your data sheet folder and the column spec (which columns are PKs and which are FKs) in one place. If you have no spec, the column naming convention (*_id) alone is enough to start with.

prompt. Enter the following into your checker.

Infer FK candidates from these data sheets. If an xxx_id column points to a PK of the same name in xxx_table, treat it as an FK. Output the result as {source_sheet, source_column, target_sheet, target_column, confidence} JSON; for anything the rules cannot settle, mark confidence as low and explain why.

verify. Review the output FK map line by line — a human, no exceptions. Self-references (prev_*), mixed-in literals (like 'player'), and polymorphic references (columns that point to different sheets depending on context) are where automatic inference goes wrong most often. Run the check with the filtered map, then classify each violation from the first run, one by one, as "truly broken" or "false violation." Exclude the false ones, but record the reason in the file.

Walk through these three steps, and the Friday-evening one-character typo stops breaking Monday's build. The check catches the typo in Saturday's small-hours nightly run, and before the owner arrives at work on Monday, a collaboration-tool task is waiting for them.

Solo Scale-Down. This check is worth having even if you work alone and have no collaboration tool. Write an FK map of ten or so lines by hand and run nothing but the Python function above, and broken references get caught. Console output or a text file is plenty for notification. The point is not the notification channel — it is the flow itself: a machine catches the reference errors human eyes cannot, and puts them in front of a human.


Key Takeaways

Next Chapter Preview

10.2 The Decision Verification 3-Layer Sensor — Where Human Review Evidence Lives

At 11 p.m., the nightly job dropped a card into the collaboration tool. The title read [integrity] D17 부합 측정 미수행 7일 경과 — "D17: conformance measurement not performed, 7 days elapsed." It was an alert that a decision applied a week earlier was still sitting in the build without anyone having confirmed it "actually worked as intended." The data was fine. Sheet formats passed, FKs passed, enums passed. But the decision had not been verified.

That gap is where this chapter starts. Data can be intact while a decision is wrong, and the place to catch that wrongness is somewhere other than the data checks. Split that place into three layers, and state explicitly how far AI assists in each layer and where a human puts the stamp — that is the decision verification 3-layer sensor.


10.2.1 Even When the Data Passes, the Decision Can Be Wrong

The check cascade runs four kinds of checks in one pass — doc-audit (document consistency), data-qa (data quality), integrity, and link (broken cross-references). When all four pass, it means "the data is fine." But a wrong decision can sit on top of fine data.

A reward sheet can be formally flawless while its numbers fuel inflation; FKs can be unique while two quests occupy the same NPC at the same time; voice can be consistent while two characters' relationship settings contradict each other — every data check passes, and the decision is wrong on every count. Data checks ask "are the cells filled in"; decision checks ask "does that value agree with other decisions, other data, and actual players." In accounting terms, the former inspects voucher forms; the latter audits the financial statements for coherence.

So decision verification gets a separate sensor from data verification. Bundle them into one check and everything gets flattened into a single pass/fail line — when it fails, you can't tell whether it was a data problem or a decision problem. Separate them and responsibility becomes clear.


10.2.2 Three Layers, Three Timings, Three Kinds of Reviewers

The core of the 3-layer sensor is that it splits the verification dimensions into three. Each layer differs in what it looks at, when it runs, and how the work divides between AI and humans.

flowchart TD
    D["Decision D17
defense_factor 1000 → 1500"] --> L1 subgraph L1["Layer 1 · Decision ↔ Decision"] L1a["AI: detect scope overlap + first-pass contradiction verdict"] L1b["Human: review 'contradiction' verdicts only + sign-off"] L1a --> L1b end subgraph L2["Layer 2 · Decision ↔ Data"] L2a["AI: extract affected data + simulation + ±% rules"] L2b["Human: interpret intent-measurement gaps"] L2a --> L2b end subgraph L3["Layer 3 · Decision ↔ Players"] L3a["AI: classify feedback + sentiment analysis"] L3b["Human: cross-check 100 samples + declare final conformance"] L3a --> L3b end L1 --> L2 --> L3 --> CARD CARD["Decision card D17
3-layer verification results + review evidence"] CARD --> ATT["human_review_attestation
reviewer · timestamp · evidence attachment enforced"] style L1 fill:#e8f0ff,stroke:#4a72c0 style L2 fill:#e8f7ed,stroke:#3a9a5a style L3 fill:#fff3e0,stroke:#d08a2a style ATT fill:#fde8e8,stroke:#c04a4a

In all three layers, the final stamp is a human's — and that stamp is exactly the evidence the human_review_attestation_evidence_mandatory atom enforces. What each layer looks at, and who signs off where, comes next, layer by layer.


10.2.3 Layer 1 — AI Narrows the Field, Humans Review Only the 'Contradictions'

This layer checks whether a new decision collides with existing ones. Decision pairs grow with the square of the decision count — 200 decisions means about 20,000 pairs. No human can read them all. So the AI runs the first filter.

# decision_conflict_check.py — Layer 1 sensor
def check_new_decision(new_decision, existing_decisions):
    conflicts = []
    for existing in existing_decisions:
        if has_overlap(new_decision.scope, existing.scope):   # mechanical 1st pass: scope intersection
            verdict = llm_judge(new_decision, existing)        # AI 2nd pass: contradiction/complement/unrelated
            if verdict.label == "contradiction":
                conflicts.append({
                    "with": existing.id,
                    "label": verdict.label,
                    "reason": verdict.reason,
                    "needs_human_review": True,                # human review flag
                })
    return conflicts

has_overlap is the mechanical first filter. scope is an atom domain (combat, narrative, economy, and so on), so the intersection computation is exact. Pairs with no overlap are discarded immediately. The 20,000 pairs typically shrink to a few dozen. Only those few dozen go to the AI.

A Complete Worked Transcript — The Layer 1 Contradiction Verdict

Here I show what actually went back and forth, end to end. No summarizing, and the ambiguous output stays exactly as it was.

① The Human's Prompt (Full Text)

Judge whether the following two decisions contradict each other.

[Decision A] (2026-03-XX) Combat global cooldown unified at 0.5 seconds across all skills
[Decision B] (2026-04-XX) Healing-type skills exempted from global cooldown application

Verdict format (this format only, no other words):
- label: contradiction | complement | unrelated
- reason: one line
- confidence: 0.0~1.0

Do not guess. If the grounds are weak, lower the confidence.

(The prompt asks whether two decisions contradict each other — decision A unifies the combat global cooldown at 0.5 seconds across all skills; decision B exempts healing-type skills from the global cooldown — and demands a fixed format and nothing else: a label of 모순/보완/무관 (contradiction/complement/unrelated), a one-line reason, and a confidence from 0.0 to 1.0, with no guessing and a lower confidence when the grounds are weak.)

② Claude's Raw Output (Unedited)

- label: complement
- reason: Decision B adds an explicit healing-type exception to the 0.5-second unification rule set by decision A, narrowing the rule rather than overturning it.
- confidence: 0.74

(The raw output reads: label — complement; reason — decision B adds an explicit healing-type exception to the 0.5-second unification rule set by decision A, narrowing the rule rather than overturning it; confidence 0.74.)

③ The Human's Verification / Rejection

The verdict is "complement" (보완), confidence 0.74. Under the rules, "complement" and "unrelated" pass automatically, but anything below confidence 0.8 requires a human to at least take one look. The reviewer, team member A, reread the original text of decision B and found something — decision B carried a rider: "however, healing skills also carry their own 1.2-second cooldown." The AI had received only a one-line summary and never got this rider as input. Team member A handled the AI verdict as accepted, with reinforcement, and put on record that "the input the AI saw was incomplete."

④ The Re-Request (After Reinforcing the Input)

[Additional rider] Decision B carries the condition that "healing skills are exempt from
the global cooldown but have their own 1.2-second cooldown." Re-judge including this rider.
- label: complement
- reason: The 1.2-second self-cooldown is longer than the global 0.5 seconds, so despite the
  exemption healing skills end up with an even longer interval. It does not conflict with A's intent (spam prevention), so the complement is reinforced.
- confidence: 0.91

(The re-request adds the missing rider — healing skills are exempt from the global cooldown but carry their own 1.2-second cooldown — and asks for a new verdict. The new output: label — complement; reason — the 1.2-second self-cooldown is longer than the global 0.5 seconds, so healing skills end up with an even longer interval despite the exemption; this does not conflict with A's intent of preventing spam, so the complement relationship is reinforced; confidence 0.91.)

The verdict is still "complement," but the reasoning is now solid and confidence rose from 0.74 to 0.91. Team member A signed off here. The point is not the result but the record of the process — the AI's first verdict, the input omission a human caught, the reinforced re-request, the final review. These four steps go into the decision card's Layer 1 evidence field exactly as they happened.

This transcript has one principle. Even the AI's 'complement' and 'unrelated' verdicts never pass unconditionally. The AI wasn't wrong — the information it received was incomplete, and the one who catches that is the person who knows the decision's original text.

The check runs at three points: immediately when a new decision is added, plus an alert; on pending-atom promotion, check first, then promote; and a nightly re-check of all pairs.


10.2.4 Layer 2 — AI Does Almost Everything, Humans Interpret the Gaps

This layer measures how a decision landed in the data and whether it matches the intent. It is the easiest to automate and the most precise. If a simulator and data sheets already exist, you only lay verification rules on top.

Take decision D17 (defense_factor 1000→1500). The sensor automatically pulls the CombatBalance sheet, the auto-simulation results, and the affected character data, then compares intent (tank survivability +49%) against measurement (simulation +52%). The conformance rules are quantitative.

Measured deviation from intent Action Who
Within ±10% Conformant (auto-pass) AI
±10–25% Alert · re-review Human interprets
Over ±25% Violation · mandatory decision re-review Human decides

The human's role here is not "the AI said conformant, so pass." Interpreting the alert band and the violation band is the human's job. The D17 simulation came in at +52%, inside ±10%, an automatic pass — but the same simulation spat out one side effect: hybrid character K_021 got +28% stronger, outside the intent. It wasn't D17's direct intent, so the conformance rule doesn't catch it. The rule passes it, but to a human eye it's an incident — catching that band is why humans exist in Layer 2.

This layer's automation rate is the highest, around 95%. The remaining 5% exists precisely because of this interpretation. A number passing the rules and that number being right for the game are two different questions.


10.2.5 Layer 3 — AI Classifies, Humans Declare Final Conformance

The hardest of the three. It checks whether the decision actually worked on real players as intended. Its inputs are live metrics from 1–2 weeks after the build ships (average tank survival time, win rate in 5v5 PvP with a tank on the team) and natural-language feedback (forums and social media).

What makes this layer distinctive is that natural-language feedback becomes verification input. About 200 forum posts and about 1,500 social media posts get categorized and sentiment-scored by the AI.

[AI feedback classification — tank-related, 1 week collected]
   positive 62%   negative 23% (mostly "tanks got way too strong")   unrelated 15%

(AI feedback classification, one week of tank-related collection: positive 62%, negative 23% — mostly "tanks got way too strong" — unrelated 15%.)

Stopping here is the trap. AI sentiment classification loses accuracy when Korean and English are mixed (it wavers on whether "탱커 강해졌다 ㅋㅋ" — roughly "tanks got stronger lol" — is praise or sarcasm). So the operating rule is: every quarter, a human classifies a sample of 100 items by hand and cross-checks them against the AI results. If the discrepancy exceeds the threshold, that quarter's classification is not trusted and a human reclassifies everything.

The final conformance declaration is made by a human. For D17, the live measurement was +44% (simulation predicted +52%, an 8% error — normal range) and feedback skewed positive. The AI organized the input — "positive majority, within intended range" — and sent it up; the one who stamped it conformant was a human. Automation is about 70%, humans 30%. This layer alone can never be fully automated. A machine cannot make the final call on what players mean.


10.2.6 Human Review Evidence Is Mandatory, Not Optional

If the final stamp in all three layers is a human's, the whole system collapses unless there is evidence that the stamp was actually placed. How do you block the case where someone merely says they reviewed and didn't? On Project A, the atom human_review_attestation_evidence_mandatory enforces this.

The atom's rule is simple and uncompromising. If 'AI verdict → human review' happened on any layer of a decision card, then the reviewer's identity, the review timestamp, and review evidence (at least one of: a reinforcement memo, a rejection rationale, or a sample cross-check result) must be attached to the card. If the evidence is empty, the card cannot be promoted to "verification complete."

When evidence is missing, the integrity_check_clickup_notify atom kicks in. On detecting an integrity failure — here, "review stamp present but no evidence" — it immediately creates a card in the collaboration tool. The 11 p.m. card in this chapter's opening scene is exactly this mechanism.

These two atoms pair up to form "verification of the verification." The 3-layer sensor verifies the decision, the attestation atom verifies that a human actually performed that verification, and the notify atom catches missing evidence and reports it. However broad the AI assistance, the last cell of responsibility is filled with the name of the person who left the evidence.


10.2.7 The Decision Card — Where Verification and Evidence Meet on One Page

The unit where the three layers' results and the review evidence converge is the decision card. One card is the complete unit for one decision, and it flows into the quarterly retrospective as input. Below is the structure of the D17 card.

Decision card D17 Change: defense_factor 1000 → 1500 · Applied 2026-03-XX Layer 1 · Decision consistency ✓ No contradicting decisions · Complement relation with 7 adjacent decisions Evidence: team member A, 2026-03-XX 14:20, 1 input-omission reinforcement memo AI first verdict → human review (confidence 0.74 → 0.91 after reinforcement) Layer 2 · Data conformance ✓ Sim +52% vs intent +49% (conformant, within ±10%) ⚠ K_021 hybrid +28% outside intent — human interpretation: follow-up decision needed Layer 3 · Player conformance ✓ Live +44% vs sim +52% (8% error, normal) · Feedback skews positive Evidence: quarterly 100-sample human cross-check done, AI classification agreement 88% Overall: ✓ Conformant (K_021 side-effect follow-up decision filed in collaboration tool) attestation check: review evidence attached on all 3 layers confirmed → card promotion allowed

The red lines are the point. If the "증거:" (evidence) row of any layer is empty, the attestation atom blocks the card's promotion and the notify atom alerts the collaboration tool. When someone asks six months later, "why did we set defense_factor to 1500?", this one card answers everything — intent, measurement, live results, and the reviewer. Decision cards run on the same metadata flow as the decision-tracking atoms in Part 18.


10.2.8 Automation Rates and Adoption Order

The three layers automate to different degrees (about 80%, 95%, and 70% respectively, as the preceding sections showed). All three are partially automated and all three end with a human stamp, but overall human workload drops by more than 80%.

Adoption starts with Layer 2. If the simulator and data sheets already exist, you only add verification rules, so it pays off within 1–2 months. Then Layer 1 (small infrastructure, large effect, one more month), and last Layer 3 (the largest infrastructure, also a large effect, another 2–3 months). Trying to bolt on Layer 3 from day one and running aground is the common failure.

On the numbers: the automation rates above and the effect ratios below are based on operational observation of the author's project — author's estimates (unverified). Read them as direction and rough proportion, not as precise measurements. The ±10%/±25% conformance thresholds are actual operating rules, and the atom names (integrity_check_clickup_notify, human_review_attestation_evidence_mandatory) are real atoms.

Summarized by direction, here is what changed before and after adoption. Decision-contradiction incidents per quarter went from several to nearly zero; the rate of actually running the one-week conformance measurement after a decision went from a fraction to most; the rate of catching side effects before they became incidents went from under half to most. The most meaningful change is traceability — the share of decisions whose background can be retraced long afterward went from a minority to nearly all. The decision cards preserve the game's decision history.


10.2.9 Common Failures

Pattern Remedy
Running Layer 1 only (contradiction checks alone) Add Layers 2 and 3 to cover the missing dimensions
Adopting Layer 3 first Start with Layer 2 — smallest infrastructure first
Accepting the AI's 'complement/unrelated' verdicts uncritically Confidence thresholds + human sample review
Stamping the review without attaching evidence The attestation atom blocks promotion
Ignoring missing-evidence alerts Treat the notify atom's collaboration-tool card as unfinished work
Blind trust in AI classification of player feedback Quarterly human cross-check of 100 samples

Key Takeaways


Try It Yourself — Solo Scale-Down

setup. Collect your decision log in one file (decision id, scope, intent, date applied). Fix scope as an enum, like combat, narrative, economy. If you have no simulator, Layer 2 can start as "manual comparison of the related data sheets."

prompt. Each time a new decision appears, ask the AI about it against existing decisions, one pair at a time. Fix the format.

Judge whether the following two decisions contradict each other.
[Decision A] ...
[Decision B] ...
Output the format only: label(contradiction|complement|unrelated) / reason one line / confidence 0.0~1.0
No guessing. If grounds are weak, lower confidence.

(The prompt: judge whether the two decisions contradict each other; output the format only — a label (contradiction | complement | unrelated), a one-line reason, and a confidence from 0.0 to 1.0; no guessing, and lower the confidence when the grounds are weak.)

verify. For 'contradiction' verdicts and anything below confidence 0.8, reread the original decision text yourself and confirm. Once confirmed, always leave the reviewer's name, the time, and a memo (one of: reinforcement, rejection, or cross-check) on the decision card. If the evidence field is empty, do not promote that card to "verification complete" — that one line is the solo version of the attestation atom. Even running solo, leave the evidence for the you of six months from now.

10.3 The Alpha Gap Report — Gaps Classified in Natural Language, Prioritized by Humans

Monday morning, 9:12 a.m. The first check cascade of the week — the week the alpha build had just gone up — had finished. When check ran all four checkers in one pass (doc-audit, data-qa, integrity, link — see 10.2) and stopped, the number printed on the console was this: 47 violation candidates. How many of them were P0, what to look at first, and who should touch what — none of that was written anywhere in those 47 lines.

A checker only knows that something is wrong. It cannot judge whether this blocks the release or can wait until next week. The real bottleneck at the end of alpha was not a shortage of checkers — it was that a person spent the entire morning sorting the 47 lines the checkers spat out. This chapter reproduces, in full, one worked cycle in which an LLM classifies those 47 lines in natural language and a human takes that classification and assigns the priorities.


10.3.1 Check Results Are Not Decisions

In 10.1 we built some 30 verification atoms, and in 10.2 we set up a structure that filters decisions through a 3-layer sensor. What those two chapters produced is a log. A log is not a decision. Between the log and the decision lies a gap that a human used to fill by hand.

check cascade automatic · 47-line log The gap human, by hand: triage · priority Weekly decision owner · deadline · gate ← the morning vanishes in this gap Gap Report = the LLM classifies, the human prioritizes

Why this gap gets expensive at the end of alpha is simple. The checkers run dozens of times an hour, but the work of a person reading the 47 lines and deciding "q_142 is a dead end, so it blocks the release; voice_lint 412 waits for the writer's call" has to be redone every single time. Handing that triage labor to a natural-language model is the starting point of the Gap Report.


10.3.2 Worked Transcript — Handing 47 Lines to the LLM

Below is the actual session from that Monday morning, in which I pasted the raw log of the check cascade into Claude as-is and asked for a classification. I reproduce it without summarizing. The places where the model guessed wrong and where the human rejected it are left intact. This is the spine of the chapter.

① The Prompt (in Full)

Below are the violation candidates that the alpha build's weekly check cascade (doc-audit/data-qa/integrity/link) spat out, combined. Classify them for use in the weekly meeting.
Classify each item as P0 (release blocker)/P1 (review)/P2 (watch) with one line of reasoning each — if you're guessing, write "guess". Don't assert severity yourself, only "propose" it; I make the final call.
Group items that come from the same root, and recommend an owning area (level/narrative/balance/UI/data). For anything you can't judge, honestly set it aside as "unclassifiable, needs human check".

[raw log]
INTEGRITY  q_142    quest_graph: node 'end_021' has no outgoing edge (dead-end), zone=dungeon_021
INTEGRITY  q_158    quest_graph: reward node references item_id 'it_9920' not in item table
DATA-QA    reward_curve  dungeon_017 reward_exp delta +35.2% vs sibling median (threshold 25%)
DATA-QA    reward_curve  dungeon_017 reward_gold delta +31.0% vs sibling median
DOC-AUDIT  voice_lint    dialogue_id_412 contains banned term (refgame slang leak)
DOC-AUDIT  voice_lint    dialogue_id_511 tone drift: formality score 0.31 (npc 'elder' expects >0.7)
DOC-AUDIT  voice_lint    dialogue_id_512 tone drift: formality score 0.29
DOC-AUDIT  voice_lint    dialogue_id_513 tone drift: formality score 0.33
DATA-QA    loc_overflow  ui_btn_enhance EN string 24 chars > slot 18 (overflow 6)
DATA-QA    loc_overflow  ui_btn_enhance TH string overflow 11
LINK       doc_link      gdd_dungeon_021.md -> [reward_table] target moved (404)
INTEGRITY  k_skill       skill 'k_017_charge' cooldown 0.0 (suspicious, min expected 0.5)
... (47 lines total, the rest follow the same patterns — attached without omission)

② Claude's Raw Output (Not Summarized or Cleaned Up)

I have classified the raw log. The severities are my "proposals"; the final call must be made by a human.

## P0 candidates (possible release blockers — human confirmation required)

1. [q_142] quest_graph dead-end (dungeon_021)
   Reasoning: a terminal node with no outgoing edge → the player gets stuck mid-quest.
   A dead end blocks progression, so P0 candidate. (Reasoning is firm)
   Owner recommendation: level/narrative

2. [q_158] reward node references item_id 'it_9920' that does not exist
   Reasoning: NULL reference when granting the reward → possible crash or missing reward.
   P0 candidate. However, the item_id may simply not be entered yet (work in progress),
   so whether this is "incomplete data" or a "typo" needs a human check.
   Owner recommendation: data

3. [voice_lint dialogue_id_412] banned term (refgame slang leak)
   Reasoning: another game's proprietary term left in dialogue is a legal/brand risk.
   P0 candidate. (Banned terms are usually blocked unconditionally)
   Owner recommendation: narrative

## P1 candidates (decide after review)

4. [reward_curve dungeon_017] exp +35.2% / gold +31.0% (both above sibling median)
   ※ Two lines grouped as one item — reward-curve anomaly in the same dungeon.
   Reasoning: above the threshold (25%), but could be an "intended high-difficulty dungeon bonus".
   Whether it is balance intent or a mistake cannot be judged from data alone → P1.
   Owner recommendation: balance

5. [voice_lint 511·512·513] 3 tone-drift items (formality 0.29~0.33, elder NPC expects >0.7)
   ※ Three lines grouped as one item — consecutive dialogue tone drift for the same NPC 'elder'.
   Reasoning: a pattern of the same NPC slipping into casual speech in consecutive lines → likely
   one writer's output from a single work session. Recommend reviewing as a group, not individually.
   Guess: could be an intended character shift (an anger scene, etc.), so awaiting the writer's call. P1.
   Owner recommendation: narrative

6. [loc_overflow ui_btn_enhance] EN over by 6 chars / TH by 11
   ※ Two lines, one item. Multilingual overflow on the same button.
   Reasoning: truncation hurts UX but is not a crash. Needs slot widening or shorter translations. P1.
   Owner recommendation: UI

## P2 candidates (watch)

7. [doc_link gdd_dungeon_021.md -> reward_table 404]
   Reasoning: broken internal document link. No impact on the game build; a doc-consistency issue. P2.
   Owner recommendation: data (doc owner)

## Unclassifiable — needs human check

8. [k_017_charge cooldown 0.0]
   Reasoning: whether a cooldown of 0.0 is an "intentionally passive-like skill" or a
   "missing entry" cannot be judged without knowing the skill's design intent. The integrity
   checker flagged anything below 0.5 as suspicious, but I do not know this skill's design intent.
   → The balance owner must check this directly.

(The remaining 39 lines were also classified into one of the 7 patterns above — only representatives are expanded in this answer)

③ Human Verification and Rejections (Reproduced Verbatim)

This is where it starts to matter. I did not adopt the LLM's classification as-is. What follows is where I took a red pen to it myself, before the meeting.

Compressing the 47 lines into 7 groups was, by itself, a big win. If a human had triaged those 47 lines from scratch, the morning would have been gone. But of the 3 P0 candidates, a human demoted 1 (412), and a human promoted 1 P1 candidate (158). 60% of the classification was right, and a human fixed the expensive 30%. That ratio is exactly the boundary line of "the LLM processes, the human decides."

④ The Follow-Up — Sending the Human-Corrected Results Back to the Model

Good. I changed two items in your classification.
- q_158: confirmed P0 (it_9920 is a deleted item, broken reference)
- voice_lint_412: demoted to P1 (intentional archaic-expression quotation; exception added to the banned-term dictionary)
Reflect these two and render a one-page Gap Report in markdown for the weekly meeting. Order: summary→P0→P1→P2→trend.
I'll give you the trend numbers — last week: P0 5 items, P1 22 items, false positives 12%.

The model took this input and output, verbatim, the one-pager shown in the §Report Format section below. The two lines the human corrected were reflected exactly, and the trend numbers used the values the human supplied, as given (nothing was made up). This round trip is everything it takes to produce one Gap Report.


10.3.3 The Gap Triage Flow — The Boundary Between Automation and Humans

Distilled into a flow, the transcript above looks like this. The key point is that every bold branch point belongs to a human.

flowchart TD
    A[check cascade
doc-audit+data-qa+integrity+link] --> B[Raw violation log, 47 lines] B --> C{LLM first-pass triage} C -->|severity proposal| D[P0 candidates] C -->|severity proposal| E[P1 candidates] C -->|severity proposal| F[P2 candidates] C -->|cannot judge| G[Unclassifiable
needs human check] C -->|same root| H[Duplicate grouping] D --> I{Human verification} E --> I F --> I G --> I I -->|adopt| J[Severity finalized] I -->|promote/demote| K[Human revises] K --> J J --> L[Render one-page Gap Report] L --> M[Input to weekly meeting] M --> N{Integrity failure?} N -->|yes| O[integrity_check_clickup_notify
instant collaboration tool alert] N -->|no| P[Assign owner and deadline] style C fill:#e8f0fe style I fill:#fef3e8 style O fill:#fde8e8

The only box the LLM touches is the blue one. Every severity is finalized in the orange one (human verification), and from the red one an integrity failure fires straight into the collaboration tool. Checking, judging, and finalizing all belong to humans and atoms; the model handles the first-pass classification, once.


10.3.4 When Integrity Breaks, We Don't Wait for the Meeting

At the end of the triage flow sits the integrity_check_clickup_notify atom (10.1). Independently of the report-building stage, the moment an integrity check fails, this atom throws a card into the collaboration tool without waiting for the meeting. If the Gap Report is the weekly rhythm, this atom is the interrupt that breaks into that rhythm.

A violation that can break the build itself, like q_158 (a reference to a deleted item), cannot wait until the Monday meeting. The moment the cascade catches it, a card — "suspected P0: q_158 broken reference" — is auto-created in the collaboration tool and assigned to the data owner. The Gap Report is the back panel that re-collects those interrupts on a weekly cadence and shows them as a trend. Only when both layers run together do the two beats stay in time: urgent items immediately, the whole picture weekly.


10.3.5 Leaving Evidence of Human Review

The fact that a human verified the LLM's classification evaporates if it survives only as word of mouth. That is why the review step has the human_review_attestation_evidence_mandatory atom (10.2) attached — human review requires evidence.

Step ③ of the transcript above — the judgment that demoted 412 and promoted 158 — goes into the report footer as the reviewer ID, a timestamp, and a list of changed items. When someone asks next quarter, "why did 412 ship?", the record answers: "judged an intentional archaic quotation in the 2026-W21 review; exception added to the banned-term dictionary." Without this, an LLM classification is indistinguishable from automated output that was never verified.


10.3.6 The Report Format — Never More than One Page

The one-pager the model rendered as the result of follow-up ④ looks like this. The classification from the transcript above flowed straight into it.

# Alpha Gap Report — 2026-W21

## Summary
- Check cascade: 47 violation candidates → classified into 7 groups
- P0 confirmed 3 / P1 4 / P2 1 / unclassifiable 1
- Release blockers: q_142 (dead end), q_158 (broken reference)
- Human review changes: voice_412 demoted (P0→P1), q_158 promoted (P1→P0)

## P0 — Immediate Action (human-confirmed)
| ID | Violation | Area | Note |
|---|---|---|---|
| q_142 | dungeon_021 dead end | Level/Narrative | LLM·human agree |
| q_158 | reference to deleted it_9920 | Data | Promoted by human |

## P1 — Decide After Review
- reward_curve dungeon_017: exp+35%/gold+31% (balance, awaiting intent confirmation)
- voice 511·512·513: elder tone drift, group of 3 (narrative, writer's call)
- voice_412: archaic-expression quotation (narrative, banned-term exception applied)
- loc_overflow ui_btn_enhance: EN/TH truncation (UI)

## P2 — Watch
- doc_link 404 (doc consistency, no build impact)

## Unclassifiable — Needs Human Check
- k_017_charge cooldown 0.0 (balance, design intent unknown)

## Trend (vs. last week)
- P0: 3 (last week 5)
- P1: 4 groups (last week 22 items — counting method changed to grouped classification)
- False positives: human corrections 2/8 = 25% (last week 12%, ↑ — sample shrank after grouping)

---
Reviewed by: Lee Minsoo / 2026-W21 / 2 changes (evidence: §review log)

Note that the report does not hide the fact that the false-positive rate in the trend section went up to 25%. The sample shrank to 8 and a human corrected 2 of them, so arithmetically it is 25%. The report does not invent numbers to look good. Compared naively against last week's 12% it looks like a regression, but a one-line note carries the context: the classification method switched to grouping, so the sample changed. The principle of never drawing conclusions from a single week's ratio is at work right here.


10.3.7 Measurement — Where the Triage Labor Went

Here is a before-and-after comparison of introducing the worked Gap Report triage on my Project A. Among the figures below, the processing ratios and times are actual measurements pulled from meeting minutes and collaboration-tool timestamps; the checker false-positive rate has a sample that swings week to week, so I record direction only.

Item Before After Basis
First-pass triage of 47 lines Human, \~40 min 1 LLM pass + human review, \~12 min Pre-meeting work log (measured)
Check results reflected in decisions Only some Most Cross-checked against meeting minutes (measured; exact % not tallied)
Average P0 resolution time 3–5 days 1–2 days Collaboration-tool card creation→completion timestamps (measured)
Human correction rate of LLM triage 2/8 as of W21 Author's estimate (unverified, varies weekly)
Delay in noticing integrity failures Waited for the meeting Immediate (atom alert) Effect of adopting clickup_notify (direction)

There is a reason I do not write the 2/8 correction rate as a brag. It is one week's sample, and some weeks the model gets five items wrong. The certain gain is that triage labor dropped from 40 minutes to 12; the accuracy of the model's classification itself swings every week. Things got faster not because we trust the model, but because it processes the log into a form a human can verify within 12 minutes.


10.3.8 Common Failures

Pattern Remedy
A human hand-sorts the 47 lines every time Divide the labor: LLM first-pass triage → human verification
LLM severity finalized as-is Severity is a "proposal"; a human finalizes (step ③)
Same-root violations counted individually State the grouping request explicitly in the prompt
Review survives only as word of mouth Force evidence with the human_review_attestation atom
Urgent integrity failures wait until the meeting Instant alerts via the clickup_notify atom
The model makes up trend numbers Humans supply the trend; the model only renders (step ④)
The report grows long and no one reads it in the meeting Enforce one page; archive the raw log separately

Key Takeaways


Try It Yourself

setup 1. Collect the output of your check cascade (or whatever bundle of lint and integrity checkers you have) into a single file. 2. Agree as a team, one line each, on the three severity levels (P0 block / P1 review / P2 watch). 3. Build a template that records the reviewer ID and timestamp in the report footer.

prompt

Below is the weekly check output. Classify it for the weekly meeting. For severity (P0/P1/P2), attach one line of reasoning and only "propose" (write "guess" if guessing); I'll make the final call. Group same-root violations, and recommend an owning area (not names). For anything you can't judge, honestly set it aside as "unclassifiable".
[paste raw log]

verify 1. Verify every P0 candidate yourself, one by one, and record each demotion or promotion (step ③). 2. Pick just one grouped item and trace it back to confirm the items really share the same root. 3. Check in the footer that the trend numbers came from a human (that the model did not fill them in on its own).

Solo Scale-Down

If you work alone, you can skip the atoms, the collaboration tool, and the weekly meeting. Paste your checker output as text, get only the classification with the prompt above, verify just the 3 P0 candidates with your own eyes, and handle them on the spot. Delegating triage to the model and narrowing what you verify down to P0 only — that one move saves the most time at solo scale. The one-page report can be replaced by a single Notion memo.

Part 11 · Character Pet Mount

11.1 Naming Conventions and Skill-to-Art Mapping

Two days before the sprint deadline, a combat artist dropped a short video into the team messenger. The new warrior class's three-hit combo. The first and second hits had the whirl of the blade; the third made no sound at all. Silence. The artist said they had hooked up every sound; the sound designer said they had delivered every file. Neither was lying. The sound file was definitely in the repository — under the name combo3_swing_final_real.wav. The name the game code was looking for was sfx_K012_combo3_swing.wav. Not a single character overlaps.

Tracking down this silent hit swallowed that entire afternoon. This is not a problem of one clip or one sound file. As long as people are free to name things however they like, this incident is reborn dozens of times every quarter. This chapter is the story of turning that freedom into rules.

Questions This Chapter Answers - At a scale of 10,000 assets, why a name is a rule, not a freedom - What class of incidents gets closed off when the naming convention is enforced as an atom and auto-verified with lint - A worked transcript where AI drafts the animation, VFX, sound, and icon mapping attached to a single skill and a human adopts it

One line for readers outside games. A pool of 10,000 assets and fbx file-name formats may look like a games-only concern. But the one thing to take away is domain-agnostic — "the moment you name things freely, search, automation, and linking all lock up together." At scale, naming has to stop being a matter of taste and become a rule, and only names that have become rules can be found and used by code automatically — a principle that applies to any work involving documents, assets, or customer records.


11.1.1 The Scale of 10,000 Assets

Project A, which I direct, is a mobile-first MMORPG. The rough scale of its character animation assets is below. The number of player classes and enemy NPC types are actual operating figures; the clip counts and the overall estimate are the author's estimate (unverified).

Asset Count
Player character classes 6
Enemy NPC types 80–100
Average clips per character 100–150 (author's estimate)
Estimated total clips About 10,000–15,000 (author's estimate)

Ten thousand. That is ten thousand drawers. Standing in front of 10,000 unlabeled drawers asking "where was that attack motion" is betting on human memory. And that bet always loses. When you cannot find it, one of two things happens: the work takes twice as long, or — since you could not find it — the same motion gets built again. The latter is worse. The asset pool bloats, and later two subtly different copies of the same motion are floating around.

If names are a free-for-all, search is not the only thing that locks up. Automatic routing — "the code finds the animation file from the skill ID by itself" — locks up with it. If no rule can be read out of a name, the code has to carry a hand-written mapping table saying which file to use for each and every skill. And that table grows by hand every time a new character comes in.


11.1.2 The Five-Slot Naming Format — Pinning It Down as an Atom

Project A's animation file names are fixed to five slots.

<role>_<id>_<category>_<action>_<variant>.fbx

char_K001_idle_default_v1.fbx
char_K001_locomotion_walk_forward.fbx
char_K001_combat_attack_combo1_v2.fbx
char_K001_react_hit_heavy.fbx
enemy_E021_combat_skill_aoe_v1.fbx

All five slots follow fixed enums. The only slot that allows free input is id, and even that one is bound to the form [A-Z]\d{3}.

Slot Enum count Examples
role 4 char, enemy, pet, mount
id fixed format K001, E021, P003, M005
category 8 idle, locomotion, combat, react, death, social, cinematic, system
action 10–30 per category walk, run, attack, skill_aoe, hit_heavy
variant fixed format default, v1, v2, _short, _long

The point here is not the format itself but where the format gets entered. Write the naming convention on a single wiki page and it is a label nobody reads. I turned this convention into a single-source-of-truth atom named Char_Anim_Naming_Convention and made people, lint, and the LLM all look at this one atom and nothing else. The moment the format is pinned down as an atom instead of a document, naming changes character: from a "recommendation" to a "gate you must pass."

The weakness of the action slot is that its enum can grow without bound. So the standard actions are managed as a per-category dictionary.

combat:
  - attack_basic
  - attack_combo1
  - attack_combo2
  - skill_<skill_id>
  - parry
  - dodge_forward
  - dodge_back
react:
  - hit_light
  - hit_heavy
  - knockback
  - stagger
  - stun
locomotion:
  - idle
  - walk_forward
  - run_forward
  - sprint
  - jump_start
  - jump_loop
  - jump_land

Whether a new action enters the dictionary is decided by procedure. Will three or more characters use it per quarter? Is it genuinely inexpressible with existing actions? Is the category unambiguous? And most importantly — can it be absorbed as a variant? If a variant can handle it, the action list does not grow. An action dictionary staying under 100 entries is a sign of healthy operation. But I do not treat that as a hard ceiling. When a new genre or a new class comes in, 30–40 entries can be added at once. What must be blocked is not a number but unchecked proliferation.


11.1.3 Lint Blocks the Commit

Once the format is entered as an atom, you need a checker that enforces that atom automatically. A human cannot eyeball five slots every time. Below is the backbone of that lint.

# anim_naming_lint.py
import re, yaml

NAMING_PATTERN = re.compile(
    r"^(?P<role>char|enemy|pet|mount)_"
    r"(?P<id>[A-Z]\d{3})_"
    r"(?P<category>idle|locomotion|combat|react|death|social|cinematic|system)_"
    r"(?P<action>[a-z_]+?)"
    r"(?:_(?P<variant>v\d+|short|long|light|heavy|left|right|forward|back))?"
    r"\.fbx$"
)

ACTION_DICT = yaml.safe_load(open("char_anim_naming_convention.yaml"))

def check(filename):
    m = NAMING_PATTERN.match(filename)
    if not m:
        return f"Naming convention violation (5-slot format mismatch): {filename}"

    category, action = m.group("category"), m.group("action")
    # skill_<id> forms are dynamic actions, so check only the prefix
    base = "skill" if action.startswith("skill_") else action
    if base not in ACTION_DICT.get(category, []):
        return f"Outside action enum ({category}): {action}"

    return None

The moment a new fbx enters the repository, this check runs. A violation blocks the commit. What matters here is that violations are not charged to people. Instead of blaming the artist behind the silent hit, the responsibility gets shoved onto the tool: "that name should never have been committable in the first place." People make mistakes; tools block those mistakes. That is the basic posture of a naming system.

Once naming is enforced, automatic routing unlocks in return.

def play_skill_animation(character, skill_id):
    anim_path = f"char_{character.id}_combat_skill_{skill_id}.fbx"
    if not exists(anim_path):
        anim_path = f"char_{character.id}_combat_skill_default.fbx"  # fallback
    play(anim_path)

The hand-written mapping table disappears. New characters and new skills can come in, and as long as the animation files are added according to the convention, not a single line of code changes. Go back to the silent hit: if that sound file could only have entered under its convention name sfx_K012_combo3_swing.wav — then combo3_swing_final_real.wav would have bounced at the commit stage, and that afternoon would have stayed intact.

The variant slot is the safety valve that protects the action enum. Versions of the same motion (v1, v2), lengths (_short, _long), intensities (_light, _heavy), and directions (_forward, _back) are all absorbed as variants instead of letting them fork the action into micro-variants. And the game code can pick that variant by context.

def select_variant(base_action, context):
    if context.distance < 3:
        return f"{base_action}_short"
    if context.distance > 10:
        return f"{base_action}_long"
    return base_action

The naming convention becomes a branch point in the code.


11.1.4 Ten Assets per Skill — The Mapping yaml

If naming is L1, the mapping that ties skills to assets is L2. A single skill usually drags along 2–3 animations, 1–3 VFX, 2–5 sounds, and one UI icon. Call it ten assets on average. With 200 skills, that is roughly 2,000 mapping targets. Managing that scale in a human head is impossible. So each skill gets one yaml file, and that skill's assets are read from that one file only.

---
skill_id: skill_K001_combo1
description: K001 combo 1 (3-hit chain)
type: melee_combo
animations:
  - clip: char_K001_combat_attack_combo1_v2.fbx
    role: main
    bone_alignment: spine_03
vfx:
  - asset: vfx_K001_combo1_slash.vfx
    socket: weapon_tip
    timing_ms: [0, 150, 300]
  - asset: vfx_hit_blood_light.vfx
    socket: target
    timing_ms: [150]
sound:
  - asset: sfx_K001_combo1_swing.wav
    volume: 0.8
    timing_ms: 0
  - asset: sfx_hit_metal_light.wav
    volume: 0.6
    timing_ms: 150
ui_icon: icon_skill_K001_combo1.png
ui_tooltip_key: skill_K001_combo1_tooltip
verified: true
---

This one file is the entirety of one skill's assets. And every asset path inside this yaml follows the five-slot convention from 11.1. If the naming lint collapses, this mapping collapses with it. The two layers work as a pair.

Once the mapping is gathered in one place, impact tracing unlocks automatically. When you want to overhaul some VFX, there is no need to dig by hand for which skills it touches.

def find_skills_using(asset):
    affected = []
    for path in glob("skills/*.yaml"):
        skill = yaml.safe_load(open(path))
        for cat in ("vfx", "sound", "animations"):
            for entry in skill.get(cat, []):
                if entry.get("asset") == asset or entry.get("clip") == asset:
                    affected.append(skill["skill_id"])
    return affected

# find_skills_using("vfx_hit_blood_light.vfx")
# → ["skill_K001_combo1", "skill_K005_combo2", "skill_E021_attack_basic", ...]

The asset-replacement meeting gets the list of affected skills attached automatically. Before anyone can ask "what does changing this affect?", the answer is already sitting next to the meeting notes.

The mapping gets its own lint, too. Does every asset file actually exist? Is there exactly one animations.main and one ui_icon? Does every timing_ms fall within the animation's length? And — does every asset path pass the 11.1 naming convention? That last item is the nail that joins the two layers. It runs automatically at build time.


11.1.5 The Naming and Mapping Verification Flow

Here is how the naming lint and the mapping lint so far chain into a single gate.

flowchart TD
    A[New asset/skill commit] --> B{5-slot naming lint
Char_Anim_Naming atom} B -->|violation| X[Block commit
return violation message] B -->|pass| C{Mapping yaml lint} C -->|asset missing / no main·icon| X C -->|asset path violates naming convention| X C -->|pass| D[Update asset pool stats] D --> E{Was this an LLM naming/mapping candidate?} E -->|yes| F[Human adopts/rejects
reversible stage] E -->|no| G[Into the build] F -->|adopt| G F -->|reject| H[Discard candidate
reversible, zero cost] G --> I{Ordering mocap/voice recording?} I -->|yes| J[Enter irreversible stage
no undo] I -->|no| K[Assets stay reversible] classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; classDef fail fill:#fee2e2,stroke:#dc2626,color:#7f1d1d; class A,D data; class B,C,E code; class F,I human; class K pass; class X,H,J fail;

Note that the reversible/irreversible boundary sits at the end of this flow. Editing yaml, LLM candidates, keyframes — all of it is reversible. If you do not like something, discard it; the cost is close to zero. But the moment you move on to motion capture shoots, voice-actor recording sessions, or signature voice casting, things turn irreversible. Actor and studio bookings, recording booths, contracts, and market perception are on the line. That is why every naming, mapping, and persona decision has to be finished right before the irreversible stage — inside the reversible territory of yaml, LLM candidates, and keyframes.


11.1.6 Worked Transcript — Letting AI Draft a New Skill Mapping

That was the system; now for where AI comes in, shown as an actual session, as-is. This is the scene where the LLM is asked to draft the asset mapping for a new flame skill, skill_K012_flame_burst.

Prompt (in full):

Below are our project's character animation naming convention (atom),
the existing asset pool stats, and one similar skill yaml.

[Naming convention]
<role>_<id>_<category>_<action>_<variant>.fbx
role: char|enemy|pet|mount / category: idle|locomotion|combat|react|death|social|cinematic|system

[Asset pool TOP used assets]
sfx_hit_metal_light.wav (47 skills), vfx_dust_kickup.vfx (38 skills),
vfx_fire_burst_med.vfx (22 skills), sfx_fire_whoosh.wav (19 skills)

[Similar skill yaml]
skill_id: skill_K005_fire_slash / type: melee_fire
animations: [char_K005_combat_attack_fire_v1.fbx]
vfx: [vfx_K005_fire_slash.vfx, vfx_fire_burst_med.vfx]
sound: [sfx_K005_fire_swing.wav, sfx_fire_whoosh.wav]

New skill: skill_K012_flame_burst (K012 warrior, AoE flame explosion, single hit)
Draft the mapping yaml for this skill. Reuse the existing asset pool as much
as possible, and mark any assets that need new production with a [NEW] tag.

Claude's raw output (no summarizing, no cleanup):

skill_id: skill_K012_flame_burst
description: K012 warrior AoE flame explosion (single hit)
type: aoe_fire
animations:
  - clip: char_K012_combat_skill_flame_burst_v1.fbx   # [NEW] new motion needed
    role: main
    bone_alignment: spine_03
vfx:
  - asset: vfx_fire_burst_med.vfx     # reuse (pool: 22 skills)
    socket: ground
    timing_ms: [200]
  - asset: vfx_K012_flame_ring.vfx    # [NEW] new, for AoE expression
    socket: ground
    timing_ms: [200, 350]
sound:
  - asset: sfx_fire_whoosh.wav        # reuse (pool: 19 skills)
    volume: 0.8
    timing_ms: 0
  - asset: sfx_explosion_low.wav      # [NEW] presumed new explosion sound
    volume: 0.9
    timing_ms: 200
ui_icon: icon_skill_K012_flame_burst.png   # [NEW]
ui_tooltip_key: skill_K012_flame_burst_tooltip
verified: false

Note: in char_K012_combat_skill_flame_burst_v1.fbx the action is skill_flame_burst, but it needs checking whether the convention requires the skill_<skill_id> form for skill actions. The convention does not state whether skill_K012 or skill_flame_burst is correct, so this is a guess.

Human verification/rejection:

That final note hit the weak spot exactly. It was not that the LLM did not know — the convention atom had that slot written ambiguously. Project A's convention actually uses the skill_<skill_id> form, so the correct file name is char_K012_combat_skill_K012_v1.fbx. The LLM's skill_flame_burst is rejected. And sfx_explosion_low.wav was tagged [NEW], but a second look at the asset pool showed sfx_impact_heavy.wav already serving as an explosion sound in 47 skills. The pool stats given to the LLM stopped at the TOP 4, so it missed this. New production gets swapped for reuse.

Follow-up request:

Rewrite with two corrections.
1. The skill action format is skill_<skill_id>. For skill K012, the action is skill_K012.
2. The asset pool has sfx_impact_heavy.wav (47 skills, explosion/impact sound).
   Reuse it instead of producing sfx_explosion_low.wav as new.
The full pool stats are as follows. [All 38 entries attached]

In this one cycle, what the LLM did was "a plausible draft," and what the human did was "spot the ambiguity in the convention, spot the gap in the pool stats, and make the reuse call." The LLM has a tendency to stamp [NEW] on candidate assets far too easily, so the reuse judgment stays in human hands to the end. Even so, writing a yaml from a blank screen and correcting a draft you can adopt or reject are very different workloads.


11.1.7 From Conservative to Progressive — The Stage Where Humans Only Adopt

The transcript above is one scene from the progressive application. Naming and mapping operations split into two stages.

In the conservative stage, people assign names and build mappings, and automation handles only verification (lint) and tracing (find_skills_using). Most MMORPG character and asset operations today sit here. In the progressive stage, the LLM produces candidates for naming drafts, mapping drafts, and even NPC persona generation, and the decision left in human hands narrows to one: which candidate to adopt.

For the progressive stage to take hold, three things must be in place. First, a naming-convention lint engine. A naming candidate from the LLM gets adopted only if it passes the same five-slot lint as a human-written one. The rejection of the LLM's skill_flame_burst in the transcript above is this gate at work. Second, an NPC persona auto-generator. Decompose character yaml into three axes — voice_profile, anim_set, skill_set — and the LLM can take a description like "a warrior in his fifties, cautious, low voice" and propose candidates for each axis separately. Writing three axes for 100 NPCs from zero, versus picking among a few candidates per persona, are different burdens. Third, a mapping candidate generator. The reverse direction of find_skills_using — a search for "existing assets that fit this new skill" — tied to the asset pool stats, proposing reuse candidates per slot. The effect cuts both ways: it lowers new production cost and raises the reuse rate.

All three run on the same infrastructure (yaml, lint, asset pool stats). They only operate when the naming convention and the mapping yaml are aligned as a single source of truth; if the alignment collapses, there is no input to give the LLM in the first place.

It is worth noting that all three were theoretically possible even in the 2010s. They were blocked in three places. Nothing could understand in natural language what a motion was, so no five-slot candidates could be proposed; splitting voice, anim, and skill apart and bundling them back was a domain of human intuition; and finding "a VFX with a similar feel" from a text description was hard. With LLM progress since 2023, all three moved into assistable territory. A large part of the progressive character-asset vision that existed only on paper has shifted into the stage of practical application.


11.1.8 Measurement — Before and After

A before-and-after comparison of the naming and mapping rollout in Project A. Search time and onboarding duration are directions I actually felt and recorded; the ratio items are measured tallies from quarterly retrospectives. Some absolute figures are, I note, the author's estimate (unverified).

Item Before After
Motion search time (animator) 5–10 minutes 30 seconds
Duplicate production rate 12–15% 1–2%
Routing code changes per new character 50–100 lines 0 lines
Missing-asset incidents on new skills 5–8 per quarter 0–1
Unused asset buildup (share of library) About 30% About 8%
New animator onboarding 2 weeks 3 days

The last row is the quietest but biggest effect. The naming convention atom itself becomes the onboarding guide. One sentence to a new animator — "name things with these five slots, and when lint blocks you, do what lint says" — and they can work on day one.


11.1.9 Common Failures

Pattern Remedy
Naming convention lives only in a wiki doc Pin it down as a single atom + enforce with lint
Unbounded action enum growth Dictionary + procedure for additions
Commits without naming verification Block commits with automatic lint
Hard-coded mapping tables in code Naming-based automatic routing
Micro-forking actions instead of using variants Absorb into the variant slot
Asset mapping scattered across code, sheets, and docs Consolidate into one yaml file
Adopting LLM mapping candidates unverified Naming lint + human reuse judgment
Charging naming violations to people Strengthen lint; move responsibility to the tool

Key Takeaways

Try It Yourself

setup — Define animation file names with the five slots <role>_<id>_<category>_<action>_<variant>.fbx, and gather the per-category action dictionary into one yaml file. Declare that yaml your team's single source of truth.

prompt — Give the LLM "[naming convention yaml] + [asset pool stats] + [one similar skill yaml]" and ask for a mapping yaml draft for a new skill. State explicitly that it must separate reused assets from assets needing new production (the [NEW] tag).

verify — Run every asset path in the LLM's output through the naming lint (the anim_naming_lint.py above). If it fails, reject it. Among candidates that pass, a human re-checks the asset pool for every [NEW] tag and judges whether reuse is possible.

Solo Scale-Down

Next Chapter Preview

11.2 Pet and Mount Systems — From 1 Template to 50 Instances

A pet list comes up at the start of a design meeting. Twelve canines, eight felines, five birds. Nobody says, "Let's build them one at a time." Unlike characters, pets are premised from day one on stamping out 50 of them. The question is not "how do we build one species well" but "how many species will share a skeleton we build once."

Each character is a distinct being to the player, so we craft characters one at a time. Pets and mounts, by contrast, are mostly variations — the same skeleton with only the color and abilities changed — so from the very start of design we lay down a mass-production pipeline: naming conventions, templates, lint. Build one species with care and then let it get copied twelve times over, and you end up with twelve wolves differing only in color, each carrying its own separate copy of the same animation clips, and a folder bloated to 4 GB. That is not mass production; that is the result of not doing mass production. The point is not how well you build, but how little you build and how much you share.

So this chapter follows one full pass end to end: define a single canine pet template in yaml, have the AI mass-produce instances that inherit its skeleton, verify them with lint, and measure what percentage gets rejected.

11.2.1 Separating Templates from Instances

All three share a similar asset structure, but they differ in how much player attention they receive. The character is the player's own self, sharing 100% of game time. A pet is a companion kept at the player's side, sharing 50\~70% of that time; a mount is a tool brought out only for travel, staying at 10\~20%. The lower the share of attention, the less detail the player notices. Pouring the same care into a mount as into a character is like managing the desk you sit at every day and the folding chair you unfold occasionally on the same budget.

So pets and mounts run on a template-instance structure. We build one template that holds the skeleton, motions, and base abilities, then layer instances on top of it that change only the color, icon, and minor abilities. An instance shares 90% of the template's assets, so what we actually build new is only the remaining 10%. Here is the split as a diagram.

Template (1 type) pet_template_canine skeleton 4 shared anims 2 shared abilities default BT 90% of assets (built once) pet_P003 (gray wolf) override: skin=gray, icon, 1 ability pet_P004 (black wolf) override: skin=black, icon, 1 ability pet_P005 (snow wolf) … up to P012 override: skin=snow, icon, 1 ability — only 10% of assets new

Build the template block on the left once, and the instances on the right only need a color, an icon, and a single ability line swapped in. The 4 GB folder I mentioned earlier is what this looks like when the separation is skipped and 90% of the assets get copied twelve times.

11.2.2 Naming and Asset Forms — One Slot Fewer Than Characters

The naming convention for pets and mounts is the character naming from 11.1 with one slot removed. Characters use the 5-slot char_<id>_<category>_<action>_<variant>, but pets and mounts drop variant and go with 4 slots. If a variant is needed, it is folded into action.

pet_<id>_<category>_<action>.fbx
mount_<id>_<category>_<action>.fbx

Examples:
pet_P003_idle_default.fbx
pet_P003_combat_bite.fbx
mount_M005_locomotion_run.fbx

The asset-mapping yaml is also lightened from the character form by removing the vfx and sound slots. An instance that carries those slots wholesale becomes a form full of blank fields, and lint throws spurious warnings every single time.

Now for the main event. Let's define one canine template and mass-produce instances from it.

11.2.3 Worked Transcript: 1 Template → Instance Mass Production → lint → Rejection Rate

Step 1 — Writing the Template yaml by Hand

Before putting the AI on mass production, a human finalizes one template by hand. This single template becomes the quality baseline for dozens of instances, so it is not automated. Here is how I set up the canine template.

# pet_template_canine.yaml
template_id: pet_template_canine
skeleton: skel_quadruped_medium      # shared medium quadruped skeleton
shared_animations:
  - clip: pet_template_canine_idle_default.fbx
  - clip: pet_template_canine_locomotion_walk.fbx
  - clip: pet_template_canine_locomotion_run.fbx
  - clip: pet_template_canine_combat_bite.fbx
shared_abilities:
  - id: pet_template_canine_passive_speed
    description: Ally movement speed +3%
  - id: pet_template_canine_active_bite
    description: Single-target bite, cooldown 12s
bt_ref: bt_pet_canine_default        # follow + combat-assist default BT
instance_overridable:                # whitelist of fields an instance may change
  - visual_skin
  - ui_icon
  - ui_tooltip_key
  - extra_ability                    # up to 1 extra ability allowed per instance

instance_overridable is the key device here. It nails down, as a whitelist, the fields an instance is allowed to touch. If the AI tries to change the skeleton or the shared animations mid-production, it has touched a field not on this list, and lint catches it. Defining what may be changed, up front, is the seatbelt of mass production.

Step 2 — Asking the AI to Mass-Produce Instances (Full Prompt)

Below is the full prompt that mass-produced 10 instances. I am printing it as is, without summarizing. (The prompt is kept in the original Korean: it instructs the AI to generate 10 canine pet instance yaml files from the template above, requires every instance to declare template: pet_template_canine, restricts overrides to the whitelist, demands natural wolf skin variations, caps extra abilities at one per instance with no overlap against shared abilities, fixes the naming rules for ui_icon and ui_tooltip_key, assigns pet_P003 through pet_P012, and forbids any prose outside the yaml blocks.)

[Prompt]
You are an assistant that helps write pet data. Based on the template below,
generate yaml for 10 wolf-family pet instances.

[Template] pet_template_canine.yaml
(full yaml above pasted here)

[Rules]
1. Every instance must explicitly declare template: pet_template_canine.
2. overrides may contain only fields from the instance_overridable whitelist.
   Never touch fields outside the whitelist (skeleton, shared_animations, etc.).
3. visual_skin must be a natural wolf variation (color, pattern, size).
4. extra_ability: at most 1 per instance. Choose 1 of passive or active.
   Its effect must not overlap with the existing shared_abilities.
5. ui_icon and ui_tooltip_key follow the naming rules matched to pet_id.
6. Assign pet_id from pet_P003 to pet_P012.

Output only the 10 yaml blocks. Do not attach explanatory prose.

Rule 2 pairs with the whitelist from Step 1. Rule 4's "must not overlap" is a constraint that stops the AI from lazily cloning abilities. Without it — as we will see shortly — the AI converges on the safest option: copy-pasting an existing ability.

Step 3 — Claude's Raw Output (No Summarizing, No Editing)

From what the AI produced, here is an excerpt with the successes, failures, and borderline cases mixed in exactly as they came out. Nothing is cleaned up.

# Output 1 — OK
pet_id: pet_P003
template: pet_template_canine
overrides:
  visual_skin: wolf_gray
  ui_icon: icon_pet_P003.png
  ui_tooltip_key: pet_P003_tooltip
  extra_ability:
    id: pet_P003_active_howl
    description: Enemy movement speed -20% for 5s, cooldown 30s

# Output 2 — OK
pet_id: pet_P004
template: pet_template_canine
overrides:
  visual_skin: wolf_black
  ui_icon: icon_pet_P004.png
  ui_tooltip_key: pet_P004_tooltip
  extra_ability:
    id: pet_P004_passive_crit
    description: Ally critical chance +2%

# Output 5 — borderline (ability effectively duplicates a shared ability)
pet_id: pet_P007
template: pet_template_canine
overrides:
  visual_skin: wolf_brown
  ui_icon: icon_pet_P007.png
  ui_tooltip_key: pet_P007_tooltip
  extra_ability:
    id: pet_P007_passive_speed_boost
    description: Ally movement speed +3%   # ← same effect as shared passive_speed

# Output 8 — FAIL (intrudes on a field outside the whitelist)
pet_id: pet_P010
template: pet_template_canine
overrides:
  visual_skin: wolf_white
  ui_icon: icon_pet_P010.png
  shared_animations:                 # ← not in the overridable whitelist
    - clip: pet_P010_combat_pounce.fbx
  extra_ability:
    id: pet_P010_active_pounce
    description: Leap attack, cooldown 20s

# Output 9 — FAIL (naming rule violation)
pet_id: P011                          # ← missing 'pet_' prefix
template: pet_template_canine
overrides:
  visual_skin: wolf_silver
  ui_icon: pet11_icon.png            # ← violates the icon_pet_P011.png rule
  ui_tooltip_key: pet_P011_tooltip
  extra_ability:
    id: pet_P011_passive_dodge
    description: Ally dodge +1%

Of the 10, six were clean — P003, P004, P005, P006, P008, P012. One was borderline on ability duplication — P007. Three failed on whitelist intrusion or naming violations — P009, P010, P011. Even with Rule 4 in place, the AI copied a shared ability in P007 (the safest choice), and even with Rule 2 in place, it touched the skeleton animations in P010. The reality is that even when constraints are spelled out, some fraction of mass-produced output leaks through. Which is why the next step exists.

Step 4 — lint Verification

Instead of a human eyeballing all 10 one by one, we run lint. The lint rules are pulled straight from the Step 1 template's whitelist and the 11.1 naming convention. There are four checks.

flowchart TD
    A[10 instance yaml files] --> B{template field
present & valid?} B -->|missing/typo| F[REJECT: template reference error] B -->|OK| C{overrides fields
within whitelist?} C -->|field outside whitelist| F2[REJECT: whitelist violation] C -->|OK| D{pet_id·ui_icon
naming rules pass?} D -->|violation| F3[REJECT: naming rule violation] D -->|OK| E{extra_ability
duplicates shared?} E -->|duplicate| W[WARN: review ability duplication] E -->|unique| P[PASS] classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; classDef fail fill:#fee2e2,stroke:#dc2626,color:#7f1d1d; class A data; class B,C,D,E code; class P pass; class F,F2,F3,W fail;

An instance that clears all four gates is PASS; one caught along the way drops to REJECT or WARN. The actual verification results, as a table:

pet_id template whitelist naming ability dup verdict
pet_P003 OK OK OK unique PASS
pet_P004 OK OK OK unique PASS
pet_P005 OK OK OK unique PASS
pet_P006 OK OK OK unique PASS
pet_P007 OK OK OK duplicate WARN
pet_P008 OK OK OK unique PASS
pet_P009 OK OK violation REJECT
pet_P010 OK intrusion REJECT
P011 OK OK violation REJECT
pet_P012 OK OK OK unique PASS

PASS 6, WARN 1, REJECT 3. The WARN can be salvaged by changing one ability line (P007); the 3 REJECTs are discarded.

Step 5 — Measuring the Rejection Rate and Re-Requesting

The rejection rate for this one cycle is REJECT 3 / 10 total = 30%. Bundling the WARN in as "needs fixing," the rework rate is 40%. This number is the health metric of the mass-production pipeline. A 30% rejection rate means that to secure 50 pets, we need to generate about 72 (50 / 0.7 ≈ 71.4). Generation is cheap, so that much overshoot is affordable. But if the rejection rate does not fall over successive rounds, that is a signal the prompt constraints are insufficient.

So we feed the rejection reasons back into the prompt. I collected the reasons behind the 3 REJECTs (missing name prefix, whitelist intrusion, icon rule violation) and added one line for each to the re-request. (The three added rules, kept in the original Korean below: pet_id must start with the pet_ prefix; ui_icon must follow the icon_<pet_id>.png format without exception; and shared_animations / skeleton / bt_ref must never appear in overrides — any behavior change must be expressed only through extra_ability.)

[Re-request additional rules]
7. pet_id must start with the 'pet_' prefix. (P011 omitted it in the previous batch)
8. ui_icon is, without exception, in the icon_<pet_id>.png format. (no variants like pet11_icon.png)
9. Never put shared_animations / skeleton / bt_ref in overrides.
   To change behavior, express it only through extra_ability. (the P010 case)

With these three lines added, the next batch of 10 came back with REJECTs down from 3 to 1. Rejection rate: 30% → 10%. This feedback — promoting rejection reasons into rules — is the mechanism that lifts mass-production quality round after round. Instead of reviewing 50 instances every time, the human's job is simply to turn each rejection reason into one line of rules.

11.2.4 Mounts — Even the Skeleton Is Shared, Almost Pure Data

Mounts are one step simpler than pets. They have no skills and no BT (Behavior Tree) — only data such as movement parameters and whether combat is allowed. So a mount instance is, in effect, a single row in a table.

# instance based on mount_template_equine.yaml
mount_id: mount_M005
template: mount_template_equine
overrides:
  visual_skin: horse_white
  movement:
    run_speed: 7.0
    sprint_speed: 12.0
  combat:
    allow_combat: false       # cannot be used in combat
    dismount_on_damage: true
  ui_icon: icon_mount_M005.png

The lint for mount mass production is even shorter. Beyond naming, template reference, and whitelist, the only added check is "are the movement parameters within the allowed range" (e.g., is sprint_speed greater than walk_speed, does it stay under the cap). It is the same pipeline built for the pets, just with fewer gates. Attaching combat features to mounts calls for caution. The moment allow_combat opens up to true, game complexity doubles, and conflict verification against the pet and character systems has to be redone from scratch.

11.2.5 Measurement — Simplification Does Not Cut into the Experience

On my Project A, I compared applying the full character pattern to pets and mounts against the simplified template-instance approach. Among the figures below, the times and asset counts are the author's estimates (unverified); the rejection rates and the asset-sharing ratio are ratios that follow the direction of actual measurement.

Item Full application Template-instance
Asset work time per pet 1\~2 weeks (author's estimate) 3\~5 days (author's estimate)
Assets in the pet library approx. 2,000 (author's estimate) approx. 600 (70% reduction)
New-asset ratio per instance 100% approx. 10%
First-batch rejection rate 30% (directional measurement)
Rejection rate after feedback 10% (directional measurement)
Player perception (pet variety) baseline nearly identical

Sample and measurement. The table above is an observation of one pet line in a single project in the author's environment (Project A) (n=1 line). The "70% reduction" and "approx. 10%" are not independent measurements but arithmetic ratios derived from the estimated asset counts in the same row (approx. 2,000 → approx. 600); since the underlying absolute values are estimates, read these percentages as estimates too. The rejection rates of 30% and 10% are a directional measurement from a single mass-production cycle, first batch through feedback, not a repeated-measurement sample. Do not cite them as evidence of savings for your team; measure your own line the same way.

The last row is the conclusion of this entire chapter. Even when mass-producing with 90% of assets shared and the rejection rate under measurement, the pet variety players felt was nearly indistinguishable from full handcrafting. The 4 GB folder from earlier is the cost paid for copying twelve sets of assets into detail players would never be able to tell apart. With mass production as the premise, what shrinks is operating cost, not the experience.

11.2.6 Operational Pitfalls

Pitfall Remedy
Porting the character system to pets and mounts as is 4-slot variation with the variant slot, vfx, and sound removed
Duplicating same-skeleton pets as independent assets 1 template + instances, with sharing enforced by the whitelist
Committing AI output without review 4 lint gates + rejection rate measurement
Rejection rate not falling round after round Promote rejection reasons into prompt rules (feedback)
Giving pets character-grade skills Cap of 1 extra_ability per instance
Giving mounts combat features Treat allow_combat with caution; expect complexity ×2

11.2.7 The AI's Place and the Human's Place

Pets and mounts touch the player experience less, so the AI gets more latitude than with characters. Matching a concept to the right template, proposing ability candidates, mass-producing instance yaml — the AI does all of this fast. But drop verification just because the latitude is wide, and the 30% rejects we saw above get mixed straight into the build. The human's place is twofold. First, finalize one template by hand to fix the quality baseline. Second, read what got filtered out and why, then refine the constraints so the next batch leaks less. The AI fills in the volume; the human holds the baseline and its corrections — that division of labor is what keeps this system running.


Key Takeaways

Next Chapter Preview


Try It Yourself

setup 1. Pick one pet family (e.g., canine), decide its common skeleton, 4 shared animations, and 2 shared abilities, and save them as pet_template_<계열>.yaml (where <계열> is the family name). 2. In the template, spell out the instance_overridable whitelist (the fields that may be changed). 3. Prepare the 4 lint gates (template reference / whitelist / naming rules / ability duplication) as a script.

prompt 4. Paste the full template yaml plus the mass-production rules (no fields outside the whitelist, no ability duplication, naming rules) and request 10 instances. 5. Lock the output format: "yaml blocks only, no explanations."

verify 6. Run lint, classify PASS / WARN / REJECT, and compute the rejection rate. 7. Collect the REJECT reasons, add one rule line per reason to the prompt, and run the next batch. Check whether the rejection rate falls.

11.2.8 Solo Scale-Down

If you are building a game solo, you can do without the lint script. Write one template yaml for one pet family by hand, then ask the AI: "5 instances from this template, changing only color, icon, and ability — never touch the skeleton or the shared animations." Skim the 5 you get back and throw out only the ones that touched the skeleton or broke the naming rules. Add one line to your next request saying why you threw them out. With just template 1.1 and that habit of feeding back the reasons you discarded, the core of this chapter works with no tooling at all.

Part 12 · Art Direction

12.1 The AI Art Asset Pipeline — Mass-Generate in Reversible Stages, Stop Before the Irreversible Gate

Primary readers: game designers and art directors who collaborate with an art team (mid-sized teams of 10–50) Scaled-down version for solo and hobbyist readers: §12.1.8, "If You're Solo, Just This Much"

I remember the day we pinned 100 sheets of AI-generated concept art to the meeting room wall. Out of the 100 sheets printed in 30 seconds, the art director picked 3, and the other 97 were thrown away on the spot. Someone called that "97% waste." But if they had been drawn by hand, an artist would have spent two weeks getting to those 3. What counted as waste had been flipped on its head.

This chapter is about turning that inversion into an operating practice. The core fits in one line. With AI art, mass-generate freely in the reversible stages (concept and texture exploration), but place a human-guarded gate before the irreversible stages (final render, motion capture, build integration). Where discarding costs nothing, throw away 99 sheets; where nothing can be undone, let not a single sheet pass unchecked. How to operate the art tools themselves is well covered in other books, so this chapter focuses only on where those tools fit safely into a game designer's pipeline.


12.1.1 The Art Pipeline Has a Line You Can Turn Back From

An art asset travels seven stages from concept to in-game. Here is the character asset line from my project (hereafter "Project A"), copied as-is. What matters is not the number of stages but the reversible/irreversible boundary that runs through the middle of them.

flowchart TB
    subgraph 가역["Reversible — discarding costs 0 (mass-generate with AI freely)"]
        direction LR
        C1["1 Concept
2D illustration"] --> C2["2 Model sheet
front/side/back"] C2 --> C3["3 3D modeling"] C3 --> C4["4 Texture
material generation"] end 가역 -.->|"Irreversible gate
must pass human review"| 비가역 subgraph 비가역["Irreversible — undoing means rework, re-recording, redeployment"] direction LR I5["5 Rigging & skinning"] --> I6["6 Animation
motion capture"] I6 --> I7["7 In-game integration
final render, live exposure"] end classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; class C1,C2,C3,C4 ai;

The four stages on the left (concept through texture) are reversible. Generate 100 concepts and discard 97, and all you lose is token cost; regenerate a texture five times, and overwriting the file is the end of it. That makes this zone the place where AI mass generation delivers the largest ROI (return on investment). The main production tools are self-hosted Stable Diffusion (SDXL) and ComfyUI. The reason is IP protection — assets never go up to an external closed service, everything runs locally, and a LoRA fine-tuned on the character plus ControlNet keep the same person consistent across repeated generations. Closed tools (Midjourney and the like) are used only sparingly for laying down early mood boards; the real mass generation, where consistency and repeat control matter, runs on SD/ComfyUI.

The three stages on the right (rigging onward) are irreversible. Motion capture ties up studio and actor schedules, and once a final render ships in a build and goes live, player memory and community reaction attach to it. Past that point, undoing costs more than making. So a human-guarded gate — a quality gate, in software terms — stands on the boundary. No matter how much AI mass-generates in the reversible zone, the only assets that cross into the irreversible are the ones that passed human review.

This one diagram is the skeleton of this chapter. The question "how much should we use AI for art" is really the question "which side of the line is this task on."


12.1.2 [Worked Transcript] One Concept Batch, from Mass Generation to Discard and Re-Request

This section shows one full cycle of concept mass generation, the first stage of the reversible zone. If you only write the abstraction — "AI generates the concepts" — you cannot tell what actually comes out and what gets discarded. Below is a faithful reproduction of a session from Project A that mass-generated concepts for a senior NPC of the scholars' guild. The prompts can be copied and used as-is; the outputs are a reconstruction of the actual session.

Step 1 — Input: State the Design Intent First

Here is where people most often get it wrong: starting the prompt with a visual description. The company feedback atom image_prompt_design_intent_first pins down the opposite principle — for image prompts too, design intent comes first. Not a string of appearance adjectives, but what function and narrative this character carries in the game, stated up front.

# concept_brief_scholar_senior.yaml — concept mass-generation input
asset_id: npc_scholar_senior_01
role: Senior of the scholars' guild — the first to observe the seal weakening
function: Main-quest-giving NPC (an informant the player must trust)
narrative_seed:
  - Has recorded the seal's veins from the bell tower for 30 years
  - Hides emotion behind numbers (scholarly_strict tone)
style_anchor: semi-realistic, painted, East Asian fantasy   # fixed from the L0 vision
forbidden: anime style · modern clothing · generic fantasy wizard robe

function and narrative_seed come before appearance. The input has to carry "why this character must look this way" before you can judge, among the mass-generated results, "why this one is better."

Step 2 — Prompt: Mass-Generate, but Enforce Format and Taboos

Based on the attached concept_brief, generate 6 proposed directions for the character concept.
This is exploratory mass generation — not finals, but candidates for the art director to pick from.

Rules:
1) Translate function and narrative_seed into visuals. No mere prettiness.
   (e.g., "hides emotion behind numbers" → how does it show in expression, props, posture)
2) Do not stray from style_anchor. The forbidden items are absolutely off-limits.
3) The 6 proposals must differ enough from one another. 6 micro-variations have no exploration value.
4) Attach a one-line intent rationale to each proposal (why this direction fits the function).
5) Separately flag any proposal you suspect 'may conflict with style_anchor'.

Output format: (proposal no. | key visual | intent rationale | conflict flag). No other explanation.

Visual description takes up almost none of this prompt. Most of it is spent locking down the output format and nailing the taboos in place. Give AI free rein and it regresses to the mean of its training data — the generic fantasy wizard. The duty to translate intent into visuals, the fixed style_anchor, the forced diversity across six proposals, and the escape hatch of "report it if you suspect a conflict" are the four bolts that block that regression.

Step 3 — Claude's Output (Reconstructed, Excerpt)

| No | Key visual                                                | Intent rationale                                    | Flag |
|----|-----------------------------------------------------------|-----------------------------------------------------|------|
| 1  | Brass magnifier permanently mounted, ink-stained fingers  | Observation is identity — the tool made flesh       |      |
| 2  | Expressionless old scholar, gauge scrolls in hand         | Hides emotion behind numbers = blank face + records |      |
| 3  | Bell tower backdrop, the seal's vein-light fading behind  | The workplace explains the person                   |      |
| 4  | Ornate star-embroidered robe, mana orb in hand            | Emphasizes majesty                                  | ★    |
| 5  | One glass eye, the mark of a lifetime of observation      | The body's price for 30 years on the job            |      |
| 6  | Young, neat scholar, looks-first emphasis                 | Likability first                                    | ★    |

The most valuable thing in this output is not the table but the two rows marked ★. On proposal 4 (the ornate wizard robe with a mana orb) and proposal 6 (looks-first), the AI itself reported "this may conflict with style_anchor/forbidden." Without the escape hatch in rule 5, those two proposals would have slipped into the table wearing the same face as the other four. Making the AI raise its own hand over the suspicious spots — that is what separates controlled mass generation from a free-for-all.

Step 4 — Verification and Rejection (the Human's Seat)

I do not take this output as-is. The art director checks the six proposals against the brief once. In this actual session, the verdicts split like this.

The two discards here are not a loss. Drawn by hand, it would have taken days to learn that those two directions were wrong; mass generation spread six proposals out at once and weeded them out within an hour.

Step 5 — Re-Request

Merge the directions of proposal 1 (magnifier made flesh) and proposal 5 (glass eye).
- Integrate the brass magnifier + one glass eye into a single figure
- Emotion suppressed (scholarly_strict): expression blank, duty spoken only through props
- Re-confirm forbidden: wizard robe, mana orb, looks-first emphasis all banned
This is the step that produces the 'single final candidate' to hand to the art director for manual finishing.

The AI answered with a single direction that merged the magnifying glass and the glass eye into one old scholar, and that one sheet went to the concept artist's desk to be finished by hand. Mass generation (6 proposals) → discard (2) → convergence (1) → human finish — one cycle closes here. What the AI produced was not the final asset but the breadth of candidates for the art director to choose from.

This one lap is the Show standard for this entire book. Unless you watch, at least once and all the way through, what the AI emits, what gets discarded, and what a human finishes, the sentence "we mass-generated concepts with AI" is hollow.


12.1.3 A High Discard Rate Is a Signal of Deep Exploration

In the session above, 2 of the 6 proposals were discarded. Across the concept line as a whole, the discards pile far higher. Of the 100 sheets pinned to the meeting room wall, 3 were adopted.

Let me handle this ratio honestly. It is a directional figure from hand-counting a few concept sessions in the early adoption period, not a precise population ratio (author's estimate, unverified — it swings widely with character personality and brief quality). So the right reading is not "exactly what percent" but the direction: compared to the hand-drawn days, we became far freer to discard.

What matters is that a 0% discard rate is not the goal. When one sheet of paper is expensive, you polish that one sheet to the end. When 100 sheets print in 30 seconds, throwing away 99 costs nothing, and the breadth of exploration widens accordingly. A rising discard rate is a signal that exploration is getting deeper. Operations that try to lower the discard rate itself — say, pressure like "if the AI made it, let's use it whenever we can" — cut down the value of exploration along with it. The reason proposals 4 and 6 in §12.1.2 could be discarded without hesitation was that discarding cost zero.


12.1.4 Texture Mass Generation — The Second Seat in the Reversible Zone

Alongside concept, the other high-ROI seat in the reversible zone is texture — the stage that generates the materials applied to 3D models. Here too, the cells where AI goes in and the cells determinism owns split cleanly.

flowchart TD
    A["UV unwrap
(human)"] --> B["Base texture
(AI-generated or painted)"] B --> C["Normal, roughness, metallic
deterministic extraction (Materialize, etc.)"] C --> D["Engine import +
lighting preview"] D --> E{"Art director review
(reversible — regenerate freely)"} E -->|"Tone mismatch"| B E -->|"Pass"| F["Asset ID & material key registration
(L3 data sheet)"] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; class C code; class B ai; class A,E human; class F data;

AI enters exactly one cell: the base texture. PBR maps such as normal, roughness, and metallic are not left to AI to produce differently every time; a deterministic extraction tool owns them. The same base has to yield the same maps for materials to stay consistent. This is the same division of labor as the city generator in §6.2, where the reward curve was never handed to AI and the rulebook held it — what determinism can guarantee goes to code; what needs exploration goes to AI.

Even the base texture is not an AI fit for every asset. Where fine detail decides the game's identity — a character's face — the human hand still comes first. That is why, when the review gate catches a "tone mismatch," the asset is sent back for regeneration rather than auto-discarded. Everything up to this point sits left of the line — the reversible zone, where any number of re-runs loses nothing.


12.1.5 The Irreversible Gate — Consistency Verification and Visual Regression

Just before the line is crossed, assets mass-generated in the reversible zone are checked for clashes with the grain of the whole game. This is a spot where human eyes alone leak, so code takes the first pass.

# visual_regression.py — detect unintended changes on asset swap (skeleton)
# Input: asset ID + before/after render captures under identical conditions
# Output: change grade (alert to the human review gate)

def compare_renders(asset_id, before_png, after_png, threshold=(1.0, 5.0)):
    diff = pixel_diff(before_png, after_png)   # normalized 0~100
    if diff > threshold[1]:
        return ("BLOCK", f"{asset_id}: major change {diff:.1f}% — no irreversible entry before review")
    elif diff > threshold[0]:
        return ("WARN",  f"{asset_id}: minor change {diff:.1f}% — intent confirmation needed")
    else:
        return ("PASS",  f"{asset_id}: no change")

These 30 lines catch the "swapped one texture and another character's shadow broke" accident before it enters the irreversible zone. The important design point is that BLOCK is not an automatic discard — it only raises an alert to the review gate. If code also kills intended changes (redesigns), the artists will be saying "turn it off" within a quarter or two. The machine picks the suspects; a human decides what crosses into the irreversible.

The other thing review catches is style consistency. AI output differs minutely on every run, so a human takes the last look at whether the mass-generated concepts and textures hold the game's grain. Only what passes this gate moves on to rigging, motion capture, and final render — the irreversible stages. Once motion is captured and shipped in a build, a consistency accident can only be fixed by rework, re-recording, and redeployment.


12.1.6 Only Decisions Cross to the Art Team — md→html→sync

Even when game designers mass-generate concepts and textures with AI, the art team that actually draws is a separate organization. The heart of this collaboration is making sure the art team never has to learn the design team's tools or conventions. Project A's art guide (96_ArtGuide/) solves this with automation.

Art decisions are written in md by the design team, _convert_md_to_html.py converts them to html, and _SyncToArtRepo.bat pushes them to a separate art repository. The art team sees only the html in that repo — they never need to know the md conventions or the design team's SVN (the pipeline diagram is in §12.2.4).

And these decision documents are split into 7 domains (00_Common and 01_Character through 07_Env), each holding its own style rules and merging at an integration gate. This is the ArtGuide's 7 areas covered in the next chapter (12.2), but here is the gist in advance — split the style rulebook into 7 drawers instead of one box, and the mass-generation prompt gets pulled from a drawer instead of being reassembled in an artist's head every time. The style_anchor and forbidden in §12.1.2 are inputs that came straight out of those drawers. The rulebook has to be separated out for mass-generated results not to regress to the generic-fantasy mean.

That said, not every game needs all 7 areas. For a casual genre, two drawers — character and environment — are plenty. Separate gradually; keep the interface narrow.


12.1.7 How to Handle Numbers and Risks Honestly

The numbers in this chapter come in only three kinds. (1) Directions and ratios — "100 generated, 3 adopted" is a directional figure from my own experience (unverified), so read it not as an absolute value but as the direction "in the reversible zone, the cost of discarding converges to zero." (2) Measured values — the visual-regression change rate (diff %), the count of consistency accidents, and the count of BLOCK events come out of visual_regression.py as numbers, so meetings can speak in numbers instead of "feelings." By contrast, "retention went up" is not decided by art alone, so I do not assert causation there.

(3) Risks stay inside the operating cost. The three risks of AI art — training-data copyright, style-consistency damage, and artists' jobs — belong inside the ROI calculation, not outside it. My policy is: AI used aggressively in the reversible zone, final assets that cross into the irreversible finished by hand, and the AI-output ratio of assets that go directly into the build held at zero as a principle. But this is one policy, not the answer — some teams use only models with explicit licenses and take AI all the way to final assets. Legal policy differs by company, and this book offers not the answer but a way to draw the line.

Of the three risks, the one most often missed is the third. Unless AI is positioned not as "a mass-production machine that replaces artists" but as "an assistant that widens exploration and strengthens artists' decision-making power," the tool will be rejected by the organization even if it succeeds on the KPIs. This is also the spot nailed down by the feedback atom design_intent_vs_automation_boundary (design intent vs. automation boundary).


12.1.8 Common Failures

Pattern Why It Fails Remedy
Putting AI concepts straight into the build as final assets The irreversible stage is passed without human review A gate before the line (§12.1.1)
Starting the prompt with appearance description Function mistranslation — regression to the pretty wizard Design intent first (§12.1.2, image_prompt_design_intent_first)
The 6 generated proposals are micro-variations No exploration value, nothing worth discarding Force diversity (§12.1.2)
Trying to lower the discard rate Cuts exploration depth along with it Read reversible-zone discards as a signal (§12.1.3)
AI-generating even the texture PBR maps Material consistency wobbles with every call Separate deterministic extraction (§12.1.4)
Swapping assets without visual regression Unintended changes leak into the irreversible The visual_regression.py gate (§12.1.5)

12.1.9 Try It Yourself — One Step You Can Take Today

If You're Solo, Just This Much: You don't need an art team or data sheets. Pick one NPC from your own game (or a game you love), write function and narrative_seed before appearance in the concept_brief format of §12.1.2, paste the 6-proposal generation prompt as-is, and run it once. Then pick the one proposal among the six that misses the intent and push back — "this is a function mistranslation; discard it and retry." Discarding in the reversible zone being exploration, not loss, will sink in through your hands.

If you're on a team, start with this one step. Draw the reversible/irreversible boundary explicitly, in one line, across your pipeline (§12.1.1). Agree on which stages are "free to discard" and where "expensive to undo" begins, and put a human review gate on that boundary. Once the line is drawn, the fight you used to restart from scratch every time — "how far should we use AI" — becomes a single ruling: "which side of the line is this task on."

Summarized as setup → prompt → verify — setup: define the reversible/irreversible boundary and the review gate in your pipeline. prompt: input the design intent first in the §12.1.2 format and mass-generate six proposals, enforcing taboos, diversity, and self-reporting. verify: in the reversible zone, pick one intent mistranslation yourself and close one full cycle of discard and re-request; before entering the irreversible, run visual_regression.py to catch unintended changes.


Key Takeaways

Next Chapter Preview


12.2 The Seven Areas of ArtGuide (Character, Animation, Monster, NPC, VFX, UI, Environment)

The Thursday integrated review. The moment we put seven new assets side by side on the same screen, we all laughed at once. The scholar character was a somber, gray-toned silhouette, and the skill VFX (visual effects) bursting right next to him was neon pink. Both were perfect decisions within their own areas. The character director had followed his own _STYLE_GUIDE.md to the letter, and the VFX artist had faithfully followed my spec: "make it highly visible." Nobody was wrong, yet placed on the same screen, two different games were at war.

That scene is both the reason ArtGuide gets split into seven areas and the reason the seven must be tied back together. The ArtGuide is the game's visual constitution. Splitting it into areas gives each area's director autonomy and speeds up decisions; failing to bind them back together through integrated reviews lets accidents like that neon pink pile up quarter after quarter. Where the designer puts a hand on this balance is the whole of this chapter.


12.2.1 One Diagram: The Actual Structure of the Seven Areas

In the design repository of Project A — the mobile-first MMORPG with an East Asian fantasy tone where I worked as design director — there is a folder named 96_ArtGuide/. The number 96 exists so that, under the repository's sort rules, the art guide lands near the end, and below it the folder splits into seven domains. This is not some abstract "art folder for the project" — what follows is the folder's actual substructure.

96_ArtGuide/ 00_Common Shared rules palette·rules 01_Character Player characters 02_Animation All animations 03_Monster Enemy NPC visuals 04_NPC Friendly NPC relations·voice 05_VFX Visual effects skills·staging 06_UI screen·HUD (9.3) 07_Environment Background·props·landmarks Each domain = one director/senior's autonomy + per-domain _STYLE_GUIDE.md (constitution) 00_Common = shared upper-level rules spanning all seven domains (color·material·period tone)

The diagram makes two points. First, the seven domains sit side by side with equal autonomy. Picture an office floor with seven studios in a row. Each room's owner holds the decision rights for that room, but when they pass each other in the hallway, the sense that they are making the same game must never be lost. Second, 00_Common sits on top of them. The shared rules all seven rooms must follow — the overall color palette, the material standards, the period tone — live here. 06_UI is the same domain as the UI collaboration standard covered in 9.1.3, so in this chapter I only draw the boundary and move on.

12.2.2 How Deeply the Designer Gets Involved Varies by Area

The designer does not intervene in all seven domains with the same intensity. The principle — the designer decides intent and narrative, art decides the visuals — applies equally to every domain, but how far intent pulls the visuals along differs by domain.

Area Designer Involvement The Line the Designer Must Not Cross
01_Character Strong Concept, personality, faction, role. Not facial proportions or brushwork
02_Animation Moderate The "type and response" of skill motions. Not frame timing
03_Monster Strong Enemy concept, faction, ecology. Not scale-pattern details
04_NPC Strong Role, relationships, voice_profile. Not costume embroidery
05_VFX Weak "Slow projectile, big explosion, purple." Not particle counts
06_UI Strong Information structure and priority (9.3). Not pixel margins
07_Environment Moderate Mood and landmark intent. Not tree polygons

The right-hand column is the real content of this table. Even in domains marked "strong," there is a line the designer must not cross. Drive the character concept hard, but the moment you touch facial proportions, the character director's autonomy collapses. And the boundary between strong and weak itself shifts with genre. In a horror game, VFX carries the fear, so designer involvement strengthens; in a casual puzzle game, character involvement actually weakens. The table above reflects Project A's genre — it is not a universal law.

12.2.3 The Domain's Constitution: _STYLE_GUIDE.md

Each domain runs on a standard bundle of documents. Here is the actual file layout of the 01_Character/ domain.

01_Character/
├── _STYLE_GUIDE.md          — overall character style (the constitution)
├── _COLOR_PALETTE.md        — color & material guide
├── _PROPORTION_REFERENCE.md — proportion & silhouette rules
├── _DO_AND_DONT.md          — allowed & forbidden
├── individual/              — per-character sheets
│   ├── K_001_director.md
│   ├── K_007_scholar.md
│   └── ...
└── _REVIEW_LOG.md           — review history

(The annotations, in Korean: _STYLE_GUIDE.md — overall character style (the constitution); _COLOR_PALETTE.md — color and material guide; _PROPORTION_REFERENCE.md — proportion and silhouette rules; _DO_AND_DONT.md — allowed and forbidden; individual/ — per-character sheets; _REVIEW_LOG.md — review history.)

_STYLE_GUIDE.md is the domain's constitution. Every individual character sheet (individual/) is a variation built on top of that constitution. If the constitution wobbles, every character under it wobbles, so this one file is the most frequently reviewed document in the domain. Its skeleton looks like this.

---
title: 01_Character Style Guide
layer: L1
---

## 1. Tone
- Korean fantasy mood, set before the 19th-century Industrial Revolution
- Realistic proportions (7~7.5 heads, no chibi deformation)

## 2. Color
- Saturation: moderate (60~70% of photoreal)
- Main palette: inherited from 00_Common
- Accent colors per character (1~2)

## 3. Costume Rules
- Costume distinguishes faction (scholar → gray + purple accent)
- Costume detail follows occupation and rank

## 4. DO
- Identifiable by silhouette alone at a 5m distance
- Express faction identity visually

## 5. DON'T
- Japanese anime style
- Non-period elements (modern clothing·props)
- Excessive saturation

(The skeleton, in Korean — 1. Tone: a Korean fantasy mood set before the 19th-century Industrial Revolution; realistic proportions (7–7.5 heads, no chibi deformation). 2. Color: moderate saturation (60–70% of photoreal); main palette: inherited from 00_Common; one or two accent colors per character. 3. Costume rules: costume distinguishes faction (scholar → gray + purple accent); detail follows occupation and rank. 4. DO: identifiable by silhouette alone at 5 m; express faction identity visually. 5. DON'T: Japanese anime style; non-period elements (modern clothing or props); excessive saturation.)

One line in there matters most. Under ## 2. 색상 (Color), the line "메인 팔레트: 00_Common 상속" — "main palette: inherited from 00_Common." It states explicitly that the character domain does not define its own colors but inherits the upper-level shared rules. This single line is the structural device that blocks the neon-pink accident from the opening. If every domain's _STYLE_GUIDE.md inherits at least its colors from 00_Common, color collisions are shut down at the constitutional level.

12.2.4 How to Bring Non-Designers into the Collaboration

Here we hit the most down-to-earth problem Project A actually faced. The art team does not read Markdown. To be precise: you must not force them to. The cost of teaching artists git diff, frontmatter, and Markdown heading hierarchies almost always exceeds whatever collaboration efficiency the teaching buys. The moment you push the design team's tools onto the art team as-is, collaboration actually slows down.

So Project A's pipeline comes down to one line: the design team decides in md, and the art team only ever sees html.

flowchart LR
    A["Design team: ArtGuide decisions
(updates _STYLE_GUIDE.md)"] --> B["_convert_md_to_html.py
(md → readable html)"] B --> C["_SyncToArtRepo.bat
(push to a separate art SVN)"] C --> D["Art team: reads html only
(zero md learning)"] D -. feedback .-> A classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef human fill:#fde68a,stroke:#b45309,color:#000; class B,C code; class A,D human;

The heart of it is two automation assets. _convert_md_to_html.py converts a domain's Markdown guides into html that artists can read comfortably in a browser. Color palettes render as actual color chips; DO/DON'T renders as visual contrast. _SyncToArtRepo.bat pushes that html not into the design repository but into a separate, art-only repository. The reason for splitting repositories is simple. If artists pull only their own repository, they are never exposed to the design team's internal md history, drafts still in progress, or other domains' decision processes. What an artist sees is only the readable end product of confirmed decisions. The cost of learning Markdown drops to zero.

This structure adds one responsibility to the designer. Whenever you update the md, you must run the convert-and-sync steps. Update without syncing, and the art team paints today's art from yesterday's decisions. If that one gap between deciding and delivering sits empty, neither autonomy nor integration means anything.

12.2.5 Image Prompts Are Decisions Too: Design Intent First

As an extension of non-designer collaboration, there is one principle a designer must keep when using generative AI at the concept stage. Project A's internal rule is named image_prompt_design_intent_first — spelled out, "in an image prompt, write the design intent first."

The common failure when exploring concepts with generated images is a prompt filled with nothing but a description of how the result should look: "an East Asian man in his fifties in a gray dopo — a traditional Korean overcoat — calm expression, photorealistic." That prompt produces a picture, but it carries nothing about why it should be that way, so the art director gets lost when trying to vary that picture. The design-intent-first principle forces an intent block above the prompt.

[Design Intent]
- Role: spiritual pillar of the scholar faction, the player's first mentor
- What must read: a "scholar, non-combatant" silhouette even at 5m
- Faction signal: scholar = gray + purple accent (inherited from 00_Common)
- Forbidden: carrying weapons, ornate armor (misread as a combat class)

[Prompt]
A male East Asian scholar in his fifties in a gray dopo, purple coat-ribbon accent,
no weapons, calm and studious expression, pre-19th-century Korean fantasy,
realistic 7.5-head proportions, moderate saturation, ...

(In Korean — [Design intent]: role — spiritual pillar of the scholar faction, the player's first mentor; what must read — a "scholar, non-combatant" silhouette even at 5 m; faction signal — scholar = gray + purple accent (inherited from 00_Common); forbidden — carrying weapons or ornate armor (would be misread as a combat class). [Prompt]: a male East Asian scholar in his fifties in a gray dopo, purple coat-ribbon accent, no weapons, calm and studious expression, pre-19th-century Korean fantasy, realistic 7.5-head proportions, moderate saturation, ...)

With the intent block sitting above the prompt, the image becomes part of a decision. When the next person generates the same character in a different pose, they are not copying the look — they are satisfying the intent all over again. Even when an image gets discarded, the discussion can be "which part of the intent failed to read?" A prompt that records only the look ends at "I just like this one better" and leaves nothing to verify against; a prompt led by intent leaves a record of what was satisfied and what was missed.

This is where tool choice locks into the intent-first principle. To repeatedly generate "the same character in a different pose" true to intent, a prompt alone is not enough. Project A's tool split follows §12.1.1 — for production runs, self-hosted Stable Diffusion (SDXL)/ComfyUI with a character LoRA (locking face and costume) and ControlNet (pose and silhouette), so the "readable at 5 m" demand in the intent block holds even as poses change, and the IP stays protected as well. Closed tools (Midjourney and the like) are used only for early mood boards.

12.2.6 Worked Transcript: One Cycle of a Character Concept Spec

To see how domain autonomy and designer involvement actually run, let me follow one cycle in 01_Character from start to finish. What follows is a reconstruction of real work in which I handed the AI the draft concept spec for the scholar character (K_007_scholar). Nothing is summarized — the awkward output and my rejection are included as they happened.

① The Prompt I Gave the AI (Full Text)

You are the character-concept assistant writer for our project. Following the
rules below, write a draft concept spec for K_007_scholar in Markdown.

Rules:
- Comply with 01_Character/_STYLE_GUIDE.md: pre-19th-century Korean fantasy,
  7~7.5-head proportions, moderate saturation, no Japanese anime style.
- Inherit colors from the 00_Common palette; the scholar faction accent is purple.
- Always include all 5 voice_profile items (tone/verbal tics/forbidden words/default emotion/forms of address).
- Decide only as far as concept, personality, faction, and role; leave visual
  details like facial proportions or brushwork blank (the character director's autonomous territory).
- Output in the individual/ character-sheet format.

(The prompt, in Korean: "You are the character-concept assistant writer for our project. Following the rules below, write a draft concept spec for K_007_scholar in Markdown." The rules: comply with 01_Character/_STYLE_GUIDE.md — pre-19th-century Korean fantasy, 7–7.5-head proportions, moderate saturation, no Japanese anime style; inherit colors from the 00_Common palette, with purple as the scholar faction's accent; always include the five voice_profile items (tone / verbal tics / forbidden words / default emotion / forms of address); decide only as far as concept, personality, faction, and role, leaving visual details like facial proportions or brushwork blank as the character director's territory; output in the individual/ character-sheet format.)

② Claude's Raw Output (Failure Included, Untouched)

# K_007 Scholar

## Role
Spiritual pillar of the scholar faction and the player's first mentor.

## Appearance
- Height 180cm, 7.5-head proportions, lean build
- White beard, deep wrinkles, piercing eyes
- Gray dopo with a purple coat ribbon
- ★ Recommended: faintly glowing golden eyes to heighten the air of mystery

## Personality
Erudite but not arrogant; strict yet warm toward his students.

## voice_profile
- Tone: slow and calm
- Verbal tic: "...geureohane", drawing out the ends of sentences
- Default emotion: equanimity
- Form of address: calls the player "young friend"

(The raw output, in Korean: a sheet titled "K_007 Scholar." Role: spiritual pillar of the scholar faction and the player's first mentor. Appearance: 180 cm, 7.5-head proportions, lean build; white beard, deep wrinkles, piercing eyes; gray dopo with a purple coat ribbon; plus a starred line — "Recommended: faintly glowing golden eyes to heighten the air of mystery." Personality: erudite but not arrogant, strict yet warm toward his students. voice_profile: tone — slow and calm; verbal tic — an archaic scholar's ending, "...geureohane," drawing out the ends of sentences; default emotion — equanimity; form of address — calls the player "young friend.")

③ My Verification and Rejection (What Was Wrong, and Why)

The output had two defects. First, the "★ 추천" (starred recommendation) line under ## 외형 (Appearance): the golden eyes. That is plain overreach. A visual detail like eye color belongs to the character director's autonomous territory, and the AI nailed it down at the concept stage. Leave it in, and the director either folds his own judgment — "design already decided this" — or ignores it and collides. One of the two. Second, voice_profile has four items, not five. "Forbidden words" is missing entirely. What a mentor character will never say is as important as his personality, and the AI dropped it.

④ The Re-Request (Correcting Exactly Two Places)

Fix exactly two things.
1. Delete the "golden eyes" recommendation line from ## Appearance. Eye color
   is the character director's decision territory. The appearance section goes only as
   far as silhouette, build, and faction color; leave specific colors and materials blank as "(director's decision)".
2. Add the missing "forbidden words" item to voice_profile. Spell out what a
   scholar-mentor would never say (profanity, vulgar jokes, modern vocabulary).

(The re-request, in Korean: "Fix exactly two things. 1. Delete the 'golden eyes' recommendation line from ## 외형 (Appearance). Eye color is the character director's decision territory. The appearance section goes only as far as silhouette, build, and faction color — leave specific colors and materials blank as '(director's decision).' 2. Add the missing 'forbidden words' item to voice_profile. Spell out what a scholar-mentor would never say: profanity, vulgar jokes, modern vocabulary.")

The lesson of this cycle is not the tool but the boundary. The AI produced a plausible draft fast, but it crossed the line the designer must not cross (visual detail) on my behalf, and it dropped something that absolutely had to be there (forbidden words). The standard of verification was not "is it well written?" but "did it respect the boundaries of domain autonomy?" As long as AI is used for character concepts, this boundary review must stay in human hands, all the way through.

12.2.7 How to Tie Cross-Area Consistency Back Together

Give seven domains autonomy and mismatches appear between them, like the neon pink in the opening. The recurring types are well known.

Mismatch Type Real Example
Character–environment tone gap The character is somber, but the background is flamboyant — they drift apart
Character–VFX color collision Gray-toned character, neon-pink skill VFX
Blurred NPC–monster boundary A friendly NPC reads as threateningly as a monster
UI–character color mismatch Cold-toned UI, warm-toned characters

The device that catches these mismatches is the weekly integrated review. The procedure is simple.

Weekly ArtGuide integrated review (Thursdays)
─────────────────────────────────
1. Randomly pull 5~10 of that week's new assets
2. Place them together on the same screen (in-game simulation)
3. Seven domain directors + game director review simultaneously
4. Mismatch found → reinforce that domain's _STYLE_GUIDE
   or the 00_Common upper-level rules

(The procedure, in Korean: Weekly ArtGuide integrated review, Thursdays — 1. randomly pull 5–10 of that week's new assets; 2. place them together on one screen, simulating in-game; 3. the seven domain directors plus the game director review simultaneously; 4. on finding a mismatch, reinforce that domain's _STYLE_GUIDE or the 00_Common upper-level rules.)

Steps 3 and 4 are the heart of it. The domain directors review simultaneously, and a discovered mismatch is not closed out by fixing the individual asset — it is fed back into the guide documents. Recolor that one neon-pink asset to gray and call it done, and the same accident returns next week. Add a rule to 00_Common instead — "skill VFX saturation stays within +20% of the character palette's saturation" — and the same accident is structurally closed off. This cycle, accumulating four times a month, is the only guardrail that keeps autonomy from hardening into silos.

12.2.8 Autonomy Is Not Distributed Responsibility but Redistributed Decision Time

Here is what the seven-area split changed, drawn from running Project A. Among the figures below, the cycle lengths and hours are the author's estimates (unverified), based on my operating experience — read them not as precise measurements but as the direction, and the rough ratio, of before versus after the split.

Item Before the Split After the Split Nature
Art decision cycle 1–2 weeks 3–5 days Author's estimate (unverified)
Cross-area consistency incidents Several per quarter Markedly fewer Direction only
Game director's art review hours Many hours a week Greatly reduced Direction only
Onboarding a new area director Months About a month Author's estimate (unverified)
Art asset discard rate High Lower Direction only

Read the table honestly, and the only thing you can assert is that every item moved in the same direction. The clearest effect is that the game director's time was reclaimed. Without autonomy, every art decision had to cross one person's desk — the game director's. Paper piles up on a single desk. Split into seven areas, the paper spread across seven desks. The total amount of paper is the same, but no desk collapses.

And the most common misunderstanding needs cutting off right here. Autonomy is not distributed responsibility. It is redistributed time. Domain directors deciding their own areas does not mean the game's overall visual responsibility shatters into seven pieces. At the integrated review, those seven come back to one table, and final visual responsibility still converges on one person. The moment autonomy becomes a pretext for dodging responsibility — the moment "that's outside my domain" becomes a catchphrase — the neon pink from the opening ships in the build as nobody's problem.

One caveat. Seven-area autonomy is a function of scale. On a small team (up to about 10 people), it is outright over-engineering. At the stage where one director wears five hats, you need not seven domain guides but a single integrated guide. Autonomy starts earning its keep only when the desks run short.

12.2.9 Common Failures and Remedies

Pattern Remedy
Every decision funnels to the game director, with no area split Adopt seven-area autonomy (but only from mid-size, 10–50 people, onward)
No domain _STYLE_GUIDE Make writing each domain's constitution a gate
No cross-area integrated verification Weekly integrated review + feeding findings back into the guides
Designer decides down to visual details Spec only intent and narrative; details belong to the director
Autonomy hardens into silos Feed back into 00_Common at the integrated review
md updated, sync skipped Make _convert_md_to_html.py_SyncToArtRepo.bat a habit
Image prompts are appearance-only Lead with the design-intent block (image_prompt_design_intent_first)

Key Takeaways

Next Chapter Preview


Try It Yourself

setup — Start with the one domain that takes the most work (usually 01_Character). In the domain folder, create _STYLE_GUIDE.md (the constitution), _COLOR_PALETTE.md, _DO_AND_DONT.md, and individual/. In the color section, be sure to enter the one line "main palette: inherited from 00_Common."

prompt — When drafting an individual asset sheet with AI, state in the prompt: (1) the rules from that domain's _STYLE_GUIDE.md, (2) "decide only as far as concept and narrative, and leave visual details blank as '(director's decision),'" and (3) the required items that must not be dropped (e.g., all five voice_profile items). For an image prompt, write the [설계 의도] (design intent) block above the appearance description first.

verify — When the output comes back, review it not for "is it well written?" but for "did it respect the boundaries of domain autonomy?" Check exactly two things: ① did the AI decide visual details the designer must not cross into, and ② are any required items missing? Point at exactly the places that are off and re-request. Every Thursday, gather 5–10 new assets on one screen and look at them together with the domain directors, simultaneously. Feed mismatches back into the guide documents (00_Common or the domain _STYLE_GUIDE), not into individual assets.

Solo Scale-Down — If you are making a game alone, do not create seven domains. Put the color palette, period tone, and DO/DON'T on a single 00_Common sheet, and divide character, environment, and VFX into sections with nothing more than headers. When you hand an asset to the AI, paste that one sheet into the prompt whole and add, "flag anywhere this guide is violated." For the integrated review: once a week, alone, a five-minute ritual of putting that week's work on one screen and looking at it. You simply have no one to share autonomy with — the skeleton of one constitution sheet plus one weekly alignment earns its keep on a team of one just the same.

12.3 Design Doc → Concept → In-Game Asset Flow

Near the end of a sprint, a concept artist dropped a character draft into the team messenger. "This is the scholar guild senior, right?" The figure on screen was a man in his thirties wearing leather armor. The game design document (GDD) — the design doc, from here on — said a woman in her forties in a gray scholar's gown. When we traced where things had gone off, it turned out the materials the concept artist had received were a two-month-old version of the design doc, and the appearance guide had changed twice in the meantime. The only person who knew it had changed was the designer.

This incident is not a technical problem. It is a flow problem. It takes 4–8 weeks on average for one page of a design doc to become an asset inside the game, and over that span, the information for a single character passes hand to hand — from the designer's head to the concept artist, to the modeler, to the animator. At every handoff the format can drift, and if a drifted handoff gets accepted anyway, the receiver fills the blanks with guesses. Two months later, the guess comes back as one line in the team messenger.

This chapter is about taking that hand-to-hand flow off one person's memory and putting it on a system.


12.3.1 Hand to Hand — Four Stages and the Handoff Points

On Project A, a character asset travels a four-stage path. What matters is not the stages themselves but the handoff points between them. Incidents do not break out inside a stage; they break out at the moment an asset is handed from one stage to the next.

Rather than explain this flow in prose, it should be drawn as a diagram — and across the 24 parts of this book, instead of drawing boxes by hand, I have asked Claude for mermaid code and rendered it. This chapter applies that very technique to its own body: the technique proving itself in its own text. Below is the unedited render of what Claude returned when I asked for "the spec→asset four-stage flow as mermaid, with the handoff gates visible."

flowchart TD
    A["Stage 1 · Design doc
character_spec.md"] -->|Spec→visual handoff| G1{Gate 1
6-item appearance check} G1 -->|Pass| B["Stage 2 · Concept art
concept_K_001_v3.png"] G1 -.->|Reject| A B -->|Visual→3D handoff| G2{Gate 2
Model sheet review} G2 -->|Pass| C["Stage 3 · 3D asset
model_K_001.fbx"] G2 -.->|Reject| B C -->|Static→dynamic handoff| G3{Gate 3
Asset lint} G3 -->|Pass| D["Stage 4 · In-game integration
Anim·VFX·Sound·Code"] G3 -.->|Reject| C D --> G4{Gate 4
Comprehensive review} G4 -->|Pass| E["Into the build"] G4 -.->|Reject| D classDef gate fill:#fde2c8,stroke:#d2691e,color:#5a2e00; classDef asset fill:#dbeafe,stroke:#2563eb,color:#0b2545; class G1,G2,G3,G4 gate; class A,B,C,D,E asset;

A gate stands at each of the three handoff points (spec→visual, visual→3D, static→dynamic). A gate is the front desk that checks the paperwork's format before passing it to the next department. If the format does not match, the item is rejected (the dotted line) and returns to the previous stage. If a non-conforming document gets accepted anyway, the next department fills the blanks with guesses. The messenger incident is what happens when Gate 1 does not exist.

The advantage of mermaid shows in this drawing. When you need to add one more gate or reorder the stages, you do not redraw boxes — you edit one line of text. Because the diagram is text, it becomes a version-control target and gets committed right alongside the design doc.


12.3.2 Stage 1 — The Design Doc Is the Root of Every Input

The flow starts from a single markdown spec. This document is the input to all three stages that follow. A blank here does not disappear; it gets pushed downstream and turns into a guess.

Below is the character_spec template as actually written. The related_atoms field connects this spec to the JIT atom system (see Part 11).

---
title: Scholar Guild Senior K_001 Character Spec
type: character_spec
layer: L2
related_atoms: [character_K_001, voice_profile_K_001]
status: draft
---

## 1. Identity
- Name: (TBD)
- Role: scholar guild senior, main NPC, recruitable as a companion
- Faction: scholar_guild
- Personality: 학자_엄격 (strict scholar), authoritative but fair

## 2. Appearance Guide
- Age: 40s
- Gender: female
- Build: slightly taller than average (around 170cm)
- Outfit: gray + purple accents, scholar's gown, glasses

## 3. Expression & Posture
- Default: calm, corners of the mouth turned down
- Anger: silence + cutting off eye contact
- Sorrow: changes the subject, minimal facial change

## 4. In-Game Role
- Appears in main quest chapters 1·5·12
- Issues 8 side quests
- Joins as a companion in chapter 7

## 5. Voice & Lines
- voice_profile: scholar_K_001
- 3 representative lines:
  - "Is that a story with no record behind it?"
  - "Do not decide on emotion. Data comes first."
  - "...So that is your choice. Understood."

## 6. References
- Images of 5 real scholars (legal review done)
- Reference the character tone of refgame

The six appearance items (age, gender, build, and outfit in section 2, plus expression and posture in section 3) become Gate 1's checklist. The two lines that caused the messenger incident — "woman in her forties / gray gown" — are exactly this section. Had Gate 1 been working, the six appearance items in the concept artist's materials would have been checked against the current design doc before the handoff.

Worked Transcript — One Round Trip to Fill the Blanks in the Design Doc

Instead of filling the whole spec by hand from zero, I lock down only the identity and role, and have Claude draft the appearance and expression sections. Below is one actual round trip, in full. Nothing is summarized.

Prompt (in full):

The identity and role of the character below are final. Based on them, draft the "2. Appearance Guide" and "3. Expression & Posture" sections of the character_spec. - Role: scholar guild senior, main NPC, recruitable as a companion - Personality: 학자_엄격 (strict scholar), authoritative but fair - Faction: scholar_guild Constraints: refgame tone (Eastern fantasy, low saturation). The appearance must be distinguishable from the other main NPCs by silhouette alone. Expressions must be restrained — anger and sorrow are never shown directly.

Claude's raw output (verbatim, untouched):

2. Appearance Guide

3. Expression & Posture

One more thing: if she "can join as a companion," you will also need combat expressions — is that outside the scope of this spec?

Human verification and veto:

I accepted half of this output. The Expression & Posture section went in almost as-is — it caught the "restrained" constraint exactly, and "sorrow replaced by turning to look out a window" is a good detail the design doc did not have. The appearance I rejected, for two reasons. First, it set the build to "slim, 165cm," while the finalized design doc said "slightly taller than average, 170cm." Claude had been given only the identity and role, so it could not know the build — it guessed. That is precisely the guess a gate exists to block. Second, the navy outfit fits the refgame tone but clashes with our game's scholar_guild colors (gray + purple). The question Claude raised ("Does the faction have a designated color?") had pointed at exactly this clash in advance.

Re-request:

Good. Expression & posture: accepted. Redo the appearance with these values fixed: build = slightly taller than average, 170cm; outfit = gray scholar's gown + purple accents (the scholar_guild colors); wears glasses. Combat expressions are out of scope for this spec — drop them.

What this single round trip teaches is that the spot where Claude filled a blank with a guess was exactly a blank in the design doc. When Claude hit a value it did not know, it split two ways. For the faction color and the combat expressions, it raised them as questions — "I don't know this" — and those questions flagged the omissions before the gate checklist did. For the build, it papered over the unknown with a plausible number and no marker. As long as the latter exists, the human pass that checks line by line against the finalized design doc cannot be skipped.


12.3.3 Stage 2 — Concept Art, and Gate 1

The finalized design doc moves to the concept artist. The flow is the same concept workflow as §12.1.2: mass-produce dozens to hundreds of images with AI, curate down to a handful, hand-polish one to three candidates, then build the model sheet (front, side, back).

What matters is Gate 1, standing at the end of this stage. Before the model sheet moves on to stage 3 (3D), the following five items are checked.

Item Pass criterion
Matches the design doc's six appearance items Outfit, build, age, gender, expression, posture match the current design doc
Silhouette distinction among main NPCs Identifiable from other characters by silhouette alone
Complies with ArtGuide 01_Character/_STYLE_GUIDE No violations of the domain style guide
No contradiction with voice_profile The visual impression does not clash with the vocal impression
Legibility at reduced size Still recognizable when shrunk to UI or minimap size

This is where the image_prompt_design_intent_first atom does its work. When the concept artist writes a prompt, they do not lead with appearance words like "gray-gowned female scholar"; they lead with the design intent from the design doc ("authoritative but fair," "a scholar who keeps emotion in check"). Mass-produce hundreds of images from appearance keywords alone and you get a pile where the gown color is right but the eyes are not a scholar's — putting intent first is how you shrink that "appearance right, impression wrong" pile in advance. The production tools are the same as §12.1.1 and §12.2.5 — self-hosted SD (SDXL)/ComfyUI with a character LoRA (locking face and outfit) plus ControlNet (locking pose and silhouette), so the face does not fall apart even across hundreds of shots of the same character in different poses.

Gate 1's first item, "matches the design doc's six appearance items," is the direct latch against the messenger incident. Because the concept draft is checked against the current design doc before it hardens into a model sheet, a drift introduced by working from a two-month-old version gets caught right here.


12.3.4 Working with Non-Designers — md for the Design Team Only, html for the Art Team Only

One operational asymmetry needs pointing out here. Every spec we have seen so far is markdown, but concept artists and 3D modelers did not join a game studio to read markdown. So Project A applies the one-way conversion pipeline from §12.2.4 ("the design team decides in md; the art team sees only html") to the spec→asset flow as well. The md decisions made by the design team are converted to html and pushed into a separate art SVN, and the art team reads only the html — their md learning cost is zero.

flowchart LR
    P["Design team
character_spec.md"] --> CV["_convert_md_to_html.py"] CV --> H["96_ArtGuide
character_spec.html"] H --> SY["_SyncToArtRepo.bat"] SY --> AR[("Art SVN
(separate repository)")] AR --> ART["Art team
Views html only"] classDef plan fill:#dcfce7,stroke:#16a34a,color:#052e16; classDef art fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef tool fill:#f3e8ff,stroke:#9333ea,color:#3b0764; class P,CV plan; class H,SY,AR,ART art; class CV,SY tool;

_convert_md_to_html.py turns the md into readable html, and _SyncToArtRepo.bat pushes the result to the art SVN, not the design SVN. The reason for keeping two repositories is the same as the PC separation principle — protecting one side's workflow from being overwritten by the other. Conversion always runs one way, design → art, and even if the art team touches the html, nothing flows back into the design md.

The destination of that conversion, 96_ArtGuide, is divided into 7 domains (00_Common, 01_Character through 07_Env). Each domain governs itself with its own _STYLE_GUIDE, while 00_Common binds the conventions shared across all domains (saturation range, naming, resolution) — the structural diagram is in §12.2.1. Gate 1's third checklist item is exactly compliance with this 01_Character/_STYLE_GUIDE.


12.3.5 Stage 3 — 3D Assets and Automated Lint

Once the model sheet moves into the 3D stage, it goes through 8 processes: high-poly modeling → retopology (game-ready low-poly) → UV unwrap → texturing → rigging and skinning → test poses → review. This is the stage where AI is weakest. 3D generative models cannot yet deliver game-quality retopology and UVs, so people and traditional tools play the lead.

Instead, this stage gets Gate 3 — automated asset lint. No one counts polygons by hand each time; the asset is checked automatically the moment it is committed.

Check Pass condition
Polygon count Standard range per character (my operating baseline: 40,000–80,000)
Texture resolution 2048×2048 standard
UV unwrap efficiency 80% or more of the area utilized
Bone count Conforms to the standard bone set
Asset naming Follows the Part 11 naming convention

When a violation is caught, the responsible 3D artist gets a notification. A check that used to depend on a sharp human eye has been moved into determinism. Items like polygon count and resolution have unambiguous right answers, so they belong neither to AI nor to humans but to a lint script.

One irreversible step appears here: the render process that bakes textures. A texture, once baked, cannot be unbaked, so Gate 3 runs once more right before the render. Stage 4's motion capture is likewise irreversible — a capture session cannot be redone short of calling back the actors and the equipment. Gates in front of irreversible steps are run more strictly than the others.


12.3.6 Stage 4 — In-Game Integration and the Comprehensive Review

Animation, VFX, sound, and code merge into the 3D asset, and the character appears inside the game for the first time. This is the stage where every discipline converges, and Gate 4 (the comprehensive review) is the final latch.

Review item Owner
Matches the design doc's intent Designer
Visual tone and consistency Art director
Animation naturalness Animation director
In-game legibility Game director
Performance (frame cost) Tech art

Five people spend 30 minutes to an hour per character. The lint at this stage is run automatically by the asset–resource mapping (Skill_Art_Resource_Mapping), checking that the resources actually wired in-game match the resources the design doc points to. In the integration stage, AI's role is confined to visual regression testing and lint automation — not deciding what to show, but the deterministic work of comparing yesterday's frame against today's, pixel by pixel, for unintended differences.


12.3.7 When a Change Touches One Stage, Everything Downstream Shakes

The messenger incident that opened this chapter was actually two incidents stacked on top of each other. One was the absence of Gate 1 (mismatched materials passed through); the other was the absence of change tracking (the appearance guide changed twice and the fact never propagated downstream). What blocks the second incident is change impact tracking.

When the materials for any stage of a character change, every artifact downstream is affected. Compute this by hand every time and you will, without fail, miss something. So we keep a tool that looks at the position in the chain and scrapes out the downstream artifacts automatically.

# spec_change_impact.py
# When any point in the chain changes, collect every downstream asset.

CHAIN = ["spec", "concept", "model", "texture", "rig", "anim", "vfx", "ingame"]

def find_downstream_artifacts(spec_id, changed_field):
    artifacts = []
    chain_position = get_chain_position(changed_field)   # e.g.: "외형.의상" → "spec"(0)
    for stage in CHAIN[chain_position + 1:]:              # everything downstream of spec
        artifacts.extend(get_artifacts(spec_id, stage))
    return artifacts

# Usage: what if K_001's outfit changes?
changed = find_downstream_artifacts("K_001", "외형.의상")
# → ["concept_K_001_v3.png", "model_K_001.fbx",
#     "texture_K_001_diffuse.png", "rig_K_001.fbx", ...]

If changed_field is "외형.의상" (appearance.outfit), the chain position is 0 (spec), and everything downstream — concept, model, texture, rig — lands on the impact list. That list goes out to the owners as automatic notifications. In the desk approval-tray metaphor: the moment tray 1 is edited, red flags pop up on trays 2 through 8, and every flagged tray goes back into the review queue. The messenger incident happened precisely because this flag did not exist — tray 1 (the design doc's appearance section) changed twice, and no flag was planted on tray 2 (the concept).


12.3.8 Measurement — The Effect of Four-Stage Standardization

Below is the before/after comparison of standardization on Project A, which I ran. The absolute times and counts are the author's estimate (unverified); what you can trust is the direction and the rough ratios.

Item Before standardization After standardization Direction
One character (design doc → in-game) 8–12 weeks 4–6 weeks Roughly halved
Guess-driven incidents between stages 10–15 per quarter 2–3 per quarter Sharp drop
Missed-change incidents 8–10 per quarter 1–2 per quarter Sharp drop
Comprehensive review time (per character) Scattered, repeated (4–6 hours total) A focused 30 minutes–1 hour Concentrated
Onboarding a new character designer About 2 months About 1 month Roughly halved

The character cycle dropped to roughly half. But do not misread this number. Standardization is not a conveyor that stamps out every character at the same speed. Main characters still get close to 8 weeks; bit parts finish in 4. What standardization did was not make the speed uniform — it made the per-stage time differentials hold steady without wobble. When a standard drifts into control, it comes back as an incident that cuts into the artists' creative time — the purpose of standardization is to eliminate guesses and omissions, not to compress time.


12.3.9 Where AI Belongs at Each Stage

Stage AI's role Strength
1. Design doc Draft assistance, asking about omissions (designer reviews) Strong
2. Concept Stable Diffusion (SDXL)/ComfyUI mass production (LoRA, ControlNet), LLM prompts Strong
3. 3D Generative models immature; people and traditional tools lead Weak
4. Integration Visual regression and lint automation Deterministic

AI is strong in stages 1 and 2, people carry stage 3, and deterministic tools own stage 4. Once this split settles in, responsibility at each stage becomes clear — where the AI's draft ends and where the human's decision begins, with no confusion at the gate.


12.3.10 Common Failures and Their Fixes

Pattern Fix
The design doc omits the six appearance/expression items Make it a mandatory stage-1 check; have the AI ask about omissions
Skipping the concept-stage gate Force the six-item appearance check before the model sheet hardens
Computing change impact by hand Automate tracking with spec_change_impact
Saving the comprehensive review for one big pass at the end Distribute gates across the stages
Building without asset lint Auto-block at Gate 3
Forcing every character into 4 weeks Keep the per-stage time differentials

The first and third rows are the direct fixes for the messenger incident that opened this chapter.


Key Takeaways

Next Chapter Preview


Try It Yourself — A Minimal spec→asset Flow

setup 1. Create one character_spec.md template (six sections — identity, the six appearance items, expression, role, voice, references — with a related_atoms field). 2. Set up an md→html conversion script (something like _convert_md_to_html.py) and share only the html with the art team. 3. Attach a gate checklist to each of the 4 handoff points (six appearance items / model sheet / asset lint / comprehensive review).

prompt

The identity and role in the character_spec below are final. Draft the "Appearance Guide" and "Expression & Posture" sections, but do not guess unknown values — mark them as questions. Constraints: refgame tone, distinguishable by silhouette alone, restrained expressions.

verify 1. Check every value the AI guessed (especially build and colors) line by line against the finalized design doc — if anything is off, reject and re-request with the values fixed. 2. Pass all five items of the Gate 1 checklist before handing off to the model sheet. 3. Deliberately change one appearance line and confirm that spec_change_impact spits out the exact downstream asset list.

Solo Scale-Down

If you work alone, the conversion pipeline, the art SVN, and the five-person review are overkill. Keep just two things. (1) One character_spec.md template — the six appearance items mandatory, blanks forbidden. (2) Every time you change the appearance, write one line at the bottom of the spec, by hand, listing the downstream files the change touches. Even without tooling, that one line blocks missed-change incidents.

Part 13 · Data Kpi

13.1 Hundreds of Free-Text Responses into Topics — AI Does the Clustering, People Do the Diagnosis

Primary audience: MMORPG designers who need to read user feedback and the metagame (mid-size teams of 10–50) Scaled-down version for solo/hobbyist readers: §13.1.8 "If You're on Your Own, Just This Much"

The morning after we shipped an update, I remember the screen showing 312 entries piled up in the free-text field of our in-game survey. They ranged from one-line fragments to five-line rants. Nobody on the design team read all 312. To be precise, nobody could. Whoever skimmed them walked into the meeting with an impression like "seems like a lot of people say enhancement is brutal" — and that impression was an illusion created by the five loudest entries. What the 312 responses actually said, nobody knew.

This chapter covers how to be able to say "what, and how many" without a person reading all 312. The core is twofold. First, hand the tedious classification — grouping hundreds of free-text responses into topics and labeling sentiment — to AI. Second, do not take AI's clusters at face value: a person catches one misclassification, rejects it, and re-requests. The general theory of FAQ and metagame analysis is in other books; this chapter focuses only on running that analysis as an AI workflow.


13.1.1 Free-Text Responses Are Something You Classify, Not Something You Read

FAQs and free-text responses are a mirror showing the gap between the game the designer intended and the game users actually experience. If the same question hits the information desk 30 times a day, you don't hire more staff — you redesign the signage. The problem is counting that "30 times." Free-text responses are not structured logs, so GROUP BY doesn't work on them. "Enhancement is too expensive" and "I can't progress because I can't earn enough currency" are the same topic but different strings. Grouping them by eye takes two or three hours for 312 entries, and the grouping criteria wobble from person to person.

This is where AI belongs. Free-text classification is (1) high-volume, (2) tedious, and (3) requires natural-language semantic judgment — that is, deterministic code can't do it, and having a person do it is expensive. But one thing must be nailed down before launch: what AI produces is a topic cluster (a hypothesis), not a confirmed diagnosis. "38% enhancement complaints" is just a result of AI attaching labels; it must not flow directly into a decision like "nerf enhancement." The principle that runs through all of Part 13 applies unchanged here — KPI definitions and final diagnosis belong to people; natural-language grouping and first-pass labeling belong to AI.

The real value of automation also lives at this point. When classification is automated, the key gain is not that the analysis itself gets faster — it's that the 312-entry signal arrives on your desk every week, already classified. The value of automation is signal exposure, not time savings (team-ops concept automation_signal_value_over_time_savings). It's the difference between letters piling up in a mailbox and mail sorted and delivered to the right department every day.


13.1.2 [Worked Transcript] 312 Free-Text Responses → Topic Clusters

Here is one full cycle of how this actually runs. Below is a faithful reproduction of a session clustering in-game survey free-text responses from my project (a mobile-first MMORPG, hereafter "Project A"). The input prompts can be copied as is; the outputs are reconstructed from the actual session.

Step 1 — Input: Throw In the Free Text As Is (No Processing)

First, extract the raw free-text responses in a machine-readable form. This just means pulling from the survey DB — you're not writing anything new. What matters is putting them in raw, with typos, profanity, and one-word answers intact — no polishing, no summarizing. Classification accuracy goes up the rawer the source text is.

# survey_freetext_2026-W21.jsonl (excerpt, 6 of 312 entries)
{"id": 0041, "text": "enhancement costs are insane -_- can't gather the currency to get to +10"}
{"id": 0088, "text": "boss patterns are fun but the rewards are way too stingy"}
{"id": 0102, "text": "guild war matchmaking takes forever, waited over 5 min"}
{"id": 0156, "text": "can't enhance without paying, you call this a game"}
{"id": 0203, "text": "new dungeon atmosphere is great, music too"}
{"id": 0274, "text": "why didn't the reward come in the mail? seems like a bug"}

Step 2 — Prompt: Ask for Clusters, but Force Categories, Evidence, and Uncertainty Reporting

Take the attached survey_freetext_2026-W21.jsonl (312 free-text survey responses), group them
into 6-10 topics, and label each response negative/neutral/positive. Don't split too finely;
each response goes in exactly one cluster. If it doesn't fit, put it in Other. For each cluster,
output a table with count, share, dominant sentiment, and one representative quote. If you're not
confident about an assignment, don't force it — pull it into an [Uncertain] list with a one-line
reason each. No commentary: just (a) cluster definitions (b) the tally table (c) the [Uncertain] list.

In this prompt, only the first two lines actually ask for classification; everything else is a constraint that boxes in the shape of the output. Give AI free rein and it will splinter the clusters into 20, or force every response into a bucket somewhere. What prevents that is the four guardrails: a cluster-count cap (6–10), single assignment, three-way sentiment, and the escape hatch — "if you're not confident, pull it into [Uncertain]." Classification accuracy comes from this fence.

Step 3 — Claude's Output (Reconstructed, Excerpt)

(a) Cluster definitions
C1 Enhancement/progression cost — complaints that enhancement and currency acquisition are excessive
C2 Spending pressure            — perception that non-paying/light-spending players are shut out of progression
C3 Content reward shortage      — feedback that rewards for bosses, dungeons, etc. are stingy relative to play
C4 Matchmaking/wait time        — complaints that guild war and matchmaking queues run long
C5 Content praise               — positive reactions to the new dungeon, presentation, music, etc.
C6 Bugs/missing rewards         — reports of undelivered rewards and errors
Other                           — one-offs that fit none of the six above

(b) Tally table
| Cluster | Count | Share | Dominant sentiment | Representative quote                       |
|----------|------|------|----------|--------------------------------------------|
| C1 Enhancement cost | 71 | 23%  | Negative     | "enhancement costs are insane -_- can't gather the currency to get to +10" |
| C2 Spending pressure | 58 | 19%  | Negative     | "can't enhance without paying, you call this a game"        |
| C3 Reward shortage | 49 | 16%  | Negative     | "boss patterns are fun but the rewards are way too stingy"        |
| C4 Matchmaking wait | 34 | 11%  | Negative     | "guild war matchmaking takes forever, waited over 5 min"      |
| C5 Content praise| 41 | 13%  | Positive     | "new dungeon atmosphere is great, music too"          |
| C6 Bugs/rewards | 28 | 9%   | Negative     | "why didn't the reward come in the mail? seems like a bug"                |
| Other       | 31 | 10%  | Mixed     | (many one-offs)                                 |

(c) [Uncertain] list (4 entries)
- id 0088 "boss patterns are fun but the rewards are way too stingy"
  → Assigned to C3 (reward shortage), but the first half also overlaps C5 (praise). Compound response.
- id 0156 "can't enhance without paying" → On the C1/C2 boundary. Sent to C2.
- (2 more)

The most valuable part of this output is not the tally table but the [Uncertain] list at the bottom. That is where the AI reports the uncertainty of its own assignments and hands them over to a person. A good prompt makes it possible for the AI to say "I'm not confident about this one."

Step 4 — Verification and Rejection (the Human's Seat)

This output must not go into a report as is. A person spot-checks the raw text directly. In this actual session, one entry got caught.

While unfolding the 58 entries in C2 (spending pressure) and skimming the raw text, id 0156 "can't enhance without paying, you call this a game" caught my eye. The AI had sent it to C2 (spending pressure). But the primary pain in this sentence is not "paying" — it's "can't enhance", which is C1 (enhancement cost). The user hit the enhancement wall and named spending as the cause of that wall; spending itself is not the core of the complaint. Yes, C1 and C2 are adjacent and easy to confuse — but if this counts toward C2, the "enhancement cost" signal looks smaller than 23%, and the enhancement curve that actually needs attention slips down the priority list. It's a boundary case where one misclassification can change the direction of a decision.

So I reject and re-request.

The C1 (enhancement cost) / C2 (spending pressure) boundary is getting blurry. Redraw it:
if the primary pain is the progression wall itself, it's C1; if it's the fairness complaint
of being shut out unless you pay, it's C2. id 0156 is C1 — "can't enhance" is the core.
Reassign the boundary cases by this rule and report only the counts that changed.

The AI redrew the boundary and moved 9 entries from C2 to C1. As a result, C1 went from 71 to 80 entries (26%) and C2 from 58 to 49 (16%). The picture — enhancement cost as the single largest topic — stayed the same, but its size sharpened from 23% to 26%. One round trip brings the signal's outline into focus. The reassignment count (9 entries) and the share change are values actually counted in this session (sample of 312, single week).

One thing to make clear here: the person did not reject "because the AI was wrong." The C2 assignment was defensible as an interpretation. What the person did was sharpen the cluster definition (= the KPI definition) and feed it back to the AI. The definition is the person's; the labor of re-combing 312 entries with that definition is the AI's.


13.1.3 The Pipeline — From Free-Text Responses to the Decision Gate

Run the session above automatically every week and it becomes a pipeline. Human hands touch only two places: setting sharp cluster definitions (front) and the gate connecting classification results to decisions (back). The grouping and labeling of 312 entries in between is the AI's run.

flowchart TB
    A["312 raw free-text responses
(extracted from survey DB, no polishing)"] --> B["Stage 1 AI: topic clustering
6-10 clusters + sentiment labels + [Uncertain] flags"] B --> C{"Stage 2 human verification
raw-text samples + boundary-case checks"} C -->|misclassification / fuzzy definition| D["Redefine cluster definitions
→ request AI reassignment"] D --> B C -->|pass| E["Weekly tally table
topic × count × sentiment"] E --> F{"Design decision gate
(director, designers)"} F --> G["Three-way triage: balance review
· UI/tutorial · bug fixes"] classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; class B ai; class C,D,F human; class A,E data;

The decisive design choice is that Stage 2 (human verification) does not auto-pass AI output. Make it auto-pass, and a boundary the AI drew wrong once will distort the signal in the same direction every week. The AI surfaces the suspect candidates (the [Uncertain] list), but whether to revise the cluster definitions is decided by a person. And the tally table is not a decision in itself — it is only input to the decision gate. "C1 enhancement cost 26%" is a signal that makes the director look into the enhancement curve, not an automatic nerf trigger.


13.1.4 The Metagame — Reading Free-Text Responses Against Behavior Logs

If free-text responses are "what users said," the metagame is "what users actually did." After launch, play patterns the designer never intended take root — that is the metagame. Things like the build meta (convergence on specific skill combinations), the routing meta (preferred farming routes), and the trading meta (user-agreed prices that diverge from official rates). Unlike free-text responses, these are measured quantitatively from behavior logs, and deterministic code (Python) does the tallying. This is not a place for AI.

The key is to read the two together. In the session above, C1 (enhancement cost) complaints were the largest at 26%. If, in the same week, the behavior logs show the build diversity index (concentration of top skill combinations) dropping, then two signals point the same way: "in words and in behavior, players are converging on one build and one progression path." When quantitative and qualitative agree, the decision gains conviction. Conversely, if the free-text responses are quiet but the behavior logs show convergence on a single build, that may be a danger signal of users feeling friction without saying so (= just before quiet churn).

flowchart LR
    A["Qualitative: free-text clusters
(AI clustering + human verification)"] --> C["Read together
same direction? diverging?"] B["Quantitative: behavior log tallies
(deterministic Python: build diversity · routes · prices)"] --> C C --> D["Design decision gate"] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; class B code; class A ai; class C,D human;

Here too the division of labor is clean. Behavior log tallying is done by code, not AI. Build share and trade prices are deterministic figures whose answers must not vary from call to call. AI is used only for grouping the unstructured text of free-text responses; quantitative KPIs are nailed down by code.


13.1.5 Where the Numbers in This Chapter Come From

The percentages in this chapter follow the principle of "One Promise" in the preface. The "C1 23%→26%, 9 reassigned" in §13.1.2 are values actually counted on a sample of 312 (a single week), so they should be read not as absolutes but as a direction: "enhancement cost is the single largest topic." No causal claims are made — there is no table like "we did FAQ analysis and retention went up." What this workflow can actually measure is three things: the number of misclassifications a person overturned in cluster verification (zero is a signal the verification was perfunctory), the time it took to produce the weekly tally, and whether the quantitative and qualitative signals agree.


13.1.6 Rejection and Re-Request Are Not Tool Failure — They Are the Gate's Signal

In §13.1.2, a person overturned 9 entries assigned to C2. Run verification weekly and you get zero-to-a-few such overturns every time. What matters is that zero overturns is not the goal. If verification overturns nothing, one of two things happened — the AI was perfect (rare), or the verifier stamped approval without looking at the raw text. The latter is overwhelmingly more common.

The verification gate is actually working when one or two boundary cases get caught each week and the cluster definitions sharpen a little because of them. This is the concrete form of the general principle that a person must periodically sample-check AI classification accuracy. The misclassification that scatters the same user type across different topics accumulates week after week if you trust automated classification without review.


13.1.7 Common Failures

Pattern Why it fails Remedy
Only skimming free-text responses by eye Illusion of the 5 loudest entries representing all 312 Full classification via AI clustering (§13.1.2)
Delegating wholesale: "AI, analyze the user feedback" Clusters splinter into 20, or forced assignments Cluster-count cap, single assignment, mandatory [Uncertain]
Reporting the AI tally table without verification A boundary misclassification changes the decision's direction Check raw-text samples and boundary cases directly
Wiring tally shares straight into decisions Auto-triggering "26% complaints, so nerf" The tally table is only input to the decision gate
Reading only qualitative, ignoring behavior logs Missing the quiet churn that never speaks Read quantitative (code) and qualitative (AI) together (§13.1.4)
Having AI tally quantitative KPIs Figures shift per call, destabilizing balance Build and price tallies belong to deterministic code

The third is the one most often missed. The tally table looks clean, so you want to trust it as is. But like the single entry id 0156, one misclassification at a boundary can flip the entire priority order. Verification is not re-reading all 312 entries — it is checking only the boundary cases of the two or three largest clusters against the raw text.


13.1.8 Try It Yourself — One Step You Can Take Today

If you're on your own, just this much: You don't need a survey DB. Collect just 30–50 store reviews or community posts for your game (or a game you love) as text, paste in the prompt from §13.1.2 verbatim, and run it once. Pick one assignment among the resulting clusters that makes you go "this looks off," and push back: "this response's primary pain is a different topic — redefine and reassign." You will feel in your hands what bundle of judgments clustering really is.

If you're on a team, start with this one step. Extract one week of free-text responses as survey_freetext_YYYY-Www.jsonl with no polishing, and run it once through the prompt in §13.1.2. Then check only the boundary cases of the two largest clusters against the raw text. Sharpen the cluster definitions once, and from then on the same prompt accumulates a reproducible weekly tally automatically.


Key Takeaways

Next Chapter Preview

13.2 KPI Definition and Tracking — Humans Define, AI Diagnoses Anomaly Signals

Primary audience: live ops/data designers responsible for operational metrics, on a mid-sized (10–50 person) team Scaled-down version for solo/hobbyist readers: §13.2.8 "If you're solo, just this much"

The same scene repeated every Monday morning. The data team's daily dashboard capture went up on the meeting screen, someone said, "DAU (daily active users) looks a bit down," and someone else answered, "That's because of last week's maintenance." The numbers were right there, but the mental work of judging whether a number was an anomaly signal or noise started over from scratch every week. And the verdict differed depending on who was talking.

Let me state this chapter's conclusion up front. In KPI work, humans have exactly two jobs: defining what to adopt as KPIs, and deciding whether to promote an anomaly signal the AI raised to a confirmed diagnosis or reject it. The two tasks wedged in between — pulling numbers out of the raw logs at the same time every day, and drafting in natural language what moved against the previous week — go to deterministic code and to AI, respectively. The general theory of KPI definition (cut the list to 5–7, watch out for Goodhart's law) is well covered in other books, so this chapter focuses only on the place where those definitions run through an AI workflow.


13.2.1 KPI Definition Is the Human's Job — What Comes After Is Not

There are two judgment calls in KPI operations that only a human can make. First, what to adopt as KPIs. Second, nailing each KPI's definition down to a single sentence. Both are value judgments about the game and cannot be delegated to AI. The decision "we count Active as 5+ minutes of play" carries inside it what the game considers healthy.

The problem is that once a definition wobbles, every number built on top of it wobbles with it. If one query counts "Active User" as one login and another as 10 minutes plus one hunt, DAU diverges wholesale. That is why protecting the consistency of the definition matters more than the definition itself — it is half of operations. And consistency checking is a job for code, not for a human head (§13.2.5).

What happens after the definitions are nailed down is not the human's seat. The daily extraction that pulls numbers at the same time every day, and the first-pass write-up that scans week-over-week movements for anomaly candidates — both repeat daily, and when humans do them the standard drifts from day to day. They are exactly the kind of work to hand down to machines and models. Extraction goes to determinism (code); the first-pass diagnosis goes to AI. The human only takes the candidates the AI raises and decides confirm or reject.

Step Who Why There
KPI selection and definition Human Value judgment about the game; cannot be delegated
Daily raw extraction Code (deterministic) Same input → same number; regression-verifiable
First-pass anomaly write-up vs. previous week AI Natural-language summarization suits AI — but only up to "hypothesis"
Confirmed diagnosis; ordering segment checks Human Promotes/rejects AI hypotheses; the seat of accountability

This division of labor is the skeleton of this entire chapter. Below, I run one cycle all the way through.


13.2.2 [Worked Transcript] Daily Dashboard Raw → Automated Anomaly Write-Up

To show how this actually runs, here is one full cycle from input to human verdict. The following reconstructs, in anonymized form, a daily KPI diagnosis session from my project (a mobile-first MMORPG, hereafter "Project A"). The raw log schema, the extraction code structure, and the prompt are carried over from the real tools; the numbers are placeholder values to show the format, not measured KPIs.

Step 1 — Input: Raw Numbers from the Deterministic Extractor

First, code pulls the KPIs from the log DB at 09:00 every day. The AI does not produce these numbers — it only receives them. The extraction output is a JSON that lines the day up against the same weekday of the previous week.

// kpi_daily_2026-06-05.json — produced by extract_kpi.py (LLM input)
{
  "date": "2026-06-05",
  "compare_to": "2026-05-29",   // same weekday of previous week (Fri)
  "active_def": "min10_hunt1", // Active definition ID in effect
  "L0": {
    "ltv_12m_est":   {"v": 0,    "prev": 0,    "delta_pct": null},
    "d30_retention": {"v": 0,    "prev": 0,    "delta_pct": null}
  },
  "L1": {
    "dau":            {"v": 0, "prev": 0, "delta_pct": -0.0},
    "session_len_min":{"v": 0, "prev": 0, "delta_pct": -0.0},
    "sessions_per_u": {"v": 0, "prev": 0, "delta_pct": 0.0},
    "d7_retention":   {"v": 0, "prev": 0, "delta_pct": 0.0}
  },
  "segments": {
    "dau_by_platform": {"ios": 0, "aos": 0},
    "dau_by_region":   {"kr": 0, "sea": 0},
    "dau_by_newbie":   {"d0_7": 0, "d8plus": 0}
  }
}

The values are left zeroed out. The structure is what matters. Each KPI carries a current value, a previous-week value, and a delta percentage, and at the very bottom rides the segment breakdown (platform, region, new vs. existing). For the AI to get past "DAU dropped" to "check which segment it dropped in," this breakdown has to be in the input.

Step 2 — The Prompt: Enforce Format, Evidence, and "No Confirmed Diagnoses"

The attached kpi_daily_2026-06-05.json is the daily KPI auto-extracted at 09:00, and
compare_to is the same weekday of the previous week. From L0~L1, pick at most 2 that
qualify as anomaly signals — ones whose week-over-week change falls outside the usual
weekday variance. If you don't know the usual variance, don't make it up; write "unknown"
and drop it. For each candidate give the metric, the change, 1~2 hypotheses for the cause,
and the segments to check — but write causes strictly as hypotheses, never as confirmed
facts. No asserting events or patches that aren't in the logs. Point only to segments
inside the attached segments object. If zero candidates, output only "0 candidates,
nothing notable."
Output 5 lines per candidate (metric/change/hypothesis/segments to check/confidence
high-mid-low), no prose.

What characterizes this prompt is that the five rule lines are almost all prohibitions. Given freedom, AI dresses a KPI diagnosis up into a plausible confirmed narrative ("DAU fell because the attendance event ended" — asserting an event that is nowhere in the logs). So I hammer in no confirmed diagnoses, hypotheses only, segments only from the input, say "unknown" when you don't know — drawing the line the AI must not cross before it starts.

Step 3 — Claude's Output (Raw, Excerpt)

Below is a reconstruction of what the LLM actually produced. I have not summarized or polished it; the spot where the third item quietly breaks a rule is left in as is.

[Anomaly signal candidates — 2]

Candidate 1
- Metric: session_len_min (average session length)
- Change: down vs. previous Friday (delta_pct negative). Largest negative change in L1.
- Hypotheses: (a) new-user (d0_7) sessions may be shorter, dragging the average down
        (b) sessions may be ending early on a specific iOS build
- Segments to check: session length split by dau_by_newbie (d0_7 vs d8plus),
                dau_by_platform (ios vs aos)
- Confidence: mid

Candidate 2
- Metric: d7_retention
- Change: slight rise vs. previous week (positive), but small in absolute terms.
- Hypothesis: possibly noise. No data on usual weekday variance, so cannot determine
  whether this is anomalous.
- Segment to check: dau_by_newbie (d0_7) cohort retention
- Confidence: low

[Note] The recent end of the attendance event appears to be the cause of the session
length drop.

The first two candidates followed the rules. Hypotheses were written as hypotheses, segments were pointed to inside the input, and candidate 2 lowered its own confidence to "low," saying it "cannot determine without usual variance data." This is what good output looks like — the AI reported its own limits.

The problem is the single [Note] line at the bottom. It asserts "the end of the attendance event" — which is nowhere in the logs — as the cause. That is a violation of rule 3. It gets caught in the next step.

Step 4 — Verification and Rejection (the Human's Seat)

Three calls to make.

First, the rule violation. The [Note] line asserted an event absent from the input JSON as if it were fact. The event calendar was not part of this input, so this is information the AI could not have known. This line is rejected.

Second, candidate 1 is accepted. The session length drop is real, and the two branches the AI proposed (new-user cohort / iOS build) can actually be checked against the input segments. Accepted — but it is still an anomaly signal, not a confirmed cause. The human's job is to run the segment queries and determine which of the two it is.

Third, candidate 2 is held. The AI itself said "cannot determine," and the absolute change is small. Until the usual weekday variance (per-weekday standard deviation) is added to the extraction code, it stays filed as noise. That is homework for the code side — when the AI reported "I don't know the usual variance," it was in fact pointing at a defect in the input data.

So I send a follow-up request.

Delete the [Note] line at the bottom — it asserts an attendance event that isn't in the
input. Keep only candidate 1, and rewrite it as a one-line action: split the session
length drop into a d0_7/d8plus × ios/aos 2x2 and "check which cell dropped the most."
No cause assertions, check actions only.

One round trip and it is done. The AI deleted the [Note] line and answered again with a single check action: "look at the d0_7 × iOS cell's session length first." That output passed the rules, and the human runs the query — and only if it confirms that the new iOS cohort cell did in fact drop the most — does the confirmed diagnosis get issued: "new-user iOS onboarding session drop-off." The diagnosis is made by a human, all the way to the end.

The point: the AI only knows "where to look." "What the cause is" gets confirmed by a human, after splitting the segments and checking. If the prompt does not enforce this boundary, the AI slides into a plausible confirmed narrative every time.


13.2.3 The KPI Pipeline — At a Glance

Pin the cycle above down as a diagram, and every daily diagnosis from then on travels the same road. You can see at a glance that human hands touch only two places, at either end (definition and confirmation).

flowchart TD
    A["KPI definition
(human: selection + one-sentence definition)"] --> B["raw log DB"] B --> C["extract_kpi.py
deterministic extraction 09:00
current · prev week · segments"] C --> D{"def_diff.py
Active definition consistency check"} D -->|mismatch alert| A D -->|match| E["AI first-pass diagnosis
anomaly signals ≤2 + hypotheses
+ segments to check"] E --> F{"human review gate
reject rule violations & assertions"} F -->|re-request| E F -->|accept| G["segment queries →
human makes confirmed diagnosis"] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; class C,D code; class E ai; class A,F,G human;

The three branches carry different colors. The blue family (extraction, definition diff) is deterministic, guaranteeing the same result for the same input. Only the single AI box in the middle is non-deterministic, which is why code holds it from both sides. The confirmed diagnosis at the far end is human. The spot where the [Note] line got caught in §13.2.2 is exactly the "F human review gate."


13.2.4 The Four Traps of KPI Definition — What Shakes Definitions

Before handing the first-pass diagnosis to AI, there are four traps lurking in the definitions humans must nail down. Miss the traps, and the input JSON of §13.2.2 itself means something different every day.

Trap 1 — The definition of Active. Whether "Active User" means one login, 5+ minutes, or 10 minutes plus one hunt splits DAU by multiples. Fix the definition as an ID (min10_hunt1) and ship it inside the input JSON (the active_def field in step 1 of §13.2.2). If this ID differs across queries, the diff in §13.2.5 catches it.

Trap 2 — When retention is measured. Whether the "7" in "7-day retention" means exactly the 7th day after signup, any day within 7 days, or the 8th day changes the value. This is an area where industry standards wobble, so the only option is to write your own definition down and keep it consistent.

Trap 3 — Outlier handling. A small number of highly active users pulls the average up. So L0–L1 are read with the median alongside the mean. A shift in the distribution often means more than a shift in the mean. Feed the AI diagnosis prompt averages only, and the AI looks at averages only — and misses the distribution shift.

Trap 4 — Measurement time. Morning, afternoon, and late-night readings differ. Operational automation standardizes on extraction at 09:00 every day, same time (step 1 of §13.2.2). If the time wobbles, the week-over-week comparison collapses.

What these four traps have in common is that it is the definition that shakes, not the value. So the most dangerous incident is not "DAU dropped" but "yesterday's DAU and today's DAU were computed under different definitions." Human eyes almost never catch it. Code does.


13.2.5 Catching Active-Definition Mismatches in Code — def_diff

The quietest KPI incident is two queries computing the same name (DAU) under different definitions. If the dashboard query counts DAU by min10_hunt1 while the marketing report query counts it by login1, two people walk into the same meeting holding different DAUs and start doubting each other. This is not something a human can catch by comparing SQL line by line, so the definition is pulled out as metadata and code diffs it.

# def_diff.py — KPI definition consistency check (skeleton)
# Premise: each query declares the Active definition ID it uses, as metadata.
#   e.g., in the dashboard.sql header:  -- @active_def: min10_hunt1

CANON = {                      # canonical definitions (nailed down once by a human)
    "DAU":          "min10_hunt1",
    "d7_retention": "signup_plus7_exact",
}

def parse_active_def(sql_path):
    # read -- @active_def: <id> from the SQL comment header
    for line in open(sql_path, encoding="utf-8"):
        if line.strip().startswith("-- @active_def:"):
            return line.split(":", 1)[1].strip()
    return None  # a missing declaration is an incident too

def diff(query_registry):
    issues = []
    for kpi, sql_path in query_registry.items():
        declared = parse_active_def(sql_path)
        canon = CANON.get(kpi)
        if declared is None:
            issues.append(f"[MISS] {kpi}: no definition declared in {sql_path}")
        elif declared != canon:
            issues.append(
                f"[DIFF] {kpi}: {sql_path} computes with '{declared}' "
                f"but canon is '{canon}'. Same name, different definition — not comparable."
            )
    return issues

These 30 lines eliminate the meeting that opens with "why is your DAU different from mine?" When code prints [DIFF] DAU: marketing_report.sql computes with 'login1' but canon is 'min10_hunt1', there is nothing left to debate. Fix the query or change the canon — one or the other. Once definitions are checked by code, you gain the guarantee that the AI diagnosis of §13.2.2 always runs on top of the same definitions. An AI diagnosis built on shaky definitions is plausible nonsense.

This check is deterministic, so it goes on CI. It runs automatically on every query commit. It is an area never delegated to AI — definition consistency is comparison, not judgment, and putting a non-deterministic model in the loop increases incidents rather than reducing them.


13.2.6 The Value of Automation Is Signal Exposure, Not Time Saved

Once this pipeline is in place, the first boast that comes to mind is "diagnosis takes less time now." The real value lies elsewhere. Among my team's operating concepts there is a one-liner named automation_signal_value_over_time_savingsthe value of automation lies not in the time saved but in the signals exposed.

Before KPI automation, a signal like a session length drop was visible only when someone happened to stare at the graph. After automation, "two anomaly signals vs. last week" lands on the desk in natural language at 09:00 every day. What shrank is analysis time; what changed is how many days it takes to become aware of the signal. What used to require a lucky glance is now forcibly exposed every day.

So this tool's success is not measured as "diagnosis takes N fewer minutes." It is measured as time to first awareness of an anomaly signal (signal → recognition). If that direction breaks — that is, if the AI summary prints "nothing notable" every day until nobody reads it — the tool has saved time while killing the signal, and within a quarter or two it becomes dead weight.


13.2.7 Where This Chapter's Numbers Come From

The numbers in this chapter follow the principle of "One Promise" from the preface. The KPI figures that appear (DAU, session length deltas) are all placeholder values to show the format, not measurements — read them as structure, not absolute values. KPI definitions (Active, retention) have no single industry-agreed standard, so the conclusion is "write your own definition down" (§13.2.4). Three things actually are measurable: the count of definition mismatches def_diff catches (target: 0), the share of AI diagnosis candidates a human rejects, and the time to anomaly-signal recognition. Conversely, I assert no causal claims like "KPI automation lifted retention."


13.2.8 Try It Yourself — One Step You Can Take Today

If you're solo, just this much: You don't need a log DB. For your own game (or a game you love), pick just 3 KPIs you would check daily and write each definition in one sentence ("Active = started at least one match"). Then jot down yesterday's and today's values by hand — two lines — paste in the prompt from §13.2.2, and have the AI "write anomaly candidates as hypotheses only, confirmed diagnoses forbidden." Find the one line where the AI quietly asserts something and push back — "that fact isn't in the logs, remove it" — and where the human's seat sits in KPI diagnosis will sink in through your hands.

If you're on a team, start with this one step. Pick 5–8 KPIs and first establish the convention of one -- @active_def: <id> line in every query's SQL header. Then put the def_diff.py skeleton from §13.2.5 (canon dict + header parsing + diff) on CI. The AI diagnosis pipeline comes after that. The definition consistency check alone blocks the quietest incident first — "your DAU and my DAU are different."


13.2.9 Common Failures

Pattern Why It Fails Remedy
A dashboard with 30 KPIs laid out Nobody can find the red, so nobody looks daily Compress to 5–8 across L0–L1
A different Active definition per query Same name, different numbers → distrust in meetings def_diff.py CI gate (§13.2.5)
Delegating "diagnose the cause" to AI wholesale Asserts events that aren't in the logs Hypotheses only; segments only from the input (§13.2.2)
Accepting AI diagnoses uncritically Plausible confirmed narratives leak into decision inputs Reject assertions at the human review gate
Feeding averages only as input Both AI and humans miss distribution shifts Include medians and segment breakdowns (§13.2.4)
Evaluating automation as "time saved" only Passes even if the summary prints nothing but "nothing notable" Measure time to signal recognition (§13.2.6)

The fourth is the one missed most often. AI summaries are smooth, and you want to take them at face value. Like the [Note] line in §13.2.2, one smooth assertion that passes without being rejected becomes a fake cause feeding into next quarter's decisions. The human's seat is not in writing the summary — it is in rejecting the summary's assertions.


Key Takeaways

Next Chapter Preview

13.3 From Anomalous Metrics to Decisions — AI Proposes Hypotheses, Humans Decide

Primary readers: data owners and directors who make quarterly decisions from KPIs (mid-size teams, 10–50 people) Scaled-down version for solo/hobbyist readers: §13.3.9, "If You're Solo, Just This Much"

One Monday morning I saw a single red line on the dashboard. Thirty-day retention had visibly broken from the previous week. Everyone in the meeting room offered a cause. Someone blamed the new hunting ground we had patched in the week before, someone blamed a competitor's new season, someone just said "seasonal factors." All of it sounded plausible. The problem was that by the end of that afternoon we had not even agreed on what to verify. Five hypotheses, and not a single segment chosen for verification.

This chapter is about how to end that kind of morning. The core fits in one line: when you see an anomalous metric, don't ask AI for a definitive diagnosis — ask it for 3–5 verifiable hypotheses. The AI does not declare "retention dropped because of X." It produces a verification design — "if X is true, this segment should look like this" — and a human makes the decision. The general theory of data-driven work is covered well enough in other books, so this chapter focuses only on running that theory as an AI workflow.


13.3.1 Humans Define the KPIs, AI Assists with Interpretation

First, I nail down the boundary. This entire chapter stands on one sentence. Humans decide what the KPIs are; AI only helps lay out, quickly, the hypotheses about why a KPI moved when it moved.

If this boundary collapses, data-driven work itself collapses. Hand KPI definition to AI and "whatever is easy to measure" becomes the KPI; hand diagnosis to AI as well and a plausible-sounding definitive sentence skips human verification and goes straight to a decision. So I open exactly one slot to AI — the stretch after an anomaly is caught and before a human decides: the stretch of "what should we suspect, and what should we check."

This division of labor shares the same spine as the earlier chapters of Part 13. Python extracts from raw logs deterministically (13.1), humans fix the KPI definitions and hierarchy (13.2), and in this chapter AI takes on only the interpretation assist when an anomaly shows up on top of that. Extraction is deterministic, definition is human, interpretation assist is AI. Keeping the three unmixed is the safety mechanism of this whole part.

My project (a mobile-first MMORPG, hereafter "Project A") has real logs underpinning this assist: under the team memory folder, _economy_log/ (token and time economy logs), _scores_latest.json (a metric score cache), and _roi_report.md (an ROI (return on investment) report). The worked transcript in this chapter takes anomaly signals extracted from these logs as its input.


13.3.2 The Decision Loop — AI Gets Exactly One Slot

Before anything else, I pin down the full loop from one anomalous metric to a decision as a diagram. In this diagram AI occupies exactly one box: "hypothesis generation." Everything before it (extraction) and after it (verification and decision) belongs to humans and code.

flowchart TB
    A["Anomaly detected
(dashboard alert / KPI threshold breach)"] A --> B["Step 1, deterministic: Python extraction
raw logs → per-segment numbers
(when, where, who dropped)"] B --> C["Step 2, AI: generate 3~5 hypotheses
definitive diagnosis banned
each hypothesis = segment to verify + expected pattern"] C --> D{"Step 3, human: prioritize hypotheses
cheapest to falsify first"} D --> E["Step 4, deterministic: Python re-extraction
precise aggregation of flagged segments only"] E --> F{"Step 5, human: decision
accept, reject, or hold each hypothesis"} F -->|Falsified| C F -->|Confirmed| G["Record decision card → apply to build"] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; class A,B,E code; class C ai; class D,F human; class G pass;

Human hands touch three places: defining what counts as anomalous (the very front, already done in 13.2), choosing which hypothesis to verify first (step 3), and making the final decision (step 5). The tedious log aggregation in between is Python's job; laying out hypotheses fast is AI's. There is no box in this loop where AI issues a definitive diagnosis. Hypotheses exist to be falsified, and a falsification loops you back to step 2.


13.3.3 [Worked Transcript] Retention Drop — Getting 3–5 Hypotheses

Here is one full cycle of how this actually runs. Below is a reconstruction of the session behind that Monday-morning retention drop. The input prompts can be copied verbatim, and the outputs faithfully reconstruct the actual session.

Step 1 — Input: Hand Over the Anomaly Signal Python Extracted, as Is

The human does not start by tossing in a feeling that "retention dropped." We toss in the per-segment numbers table that Python extracted deterministically. Nothing here is written fresh — it is extraction only, from _economy_log/event logs.

# retention_break_extract.py (skeleton) — segment breakdown of the anomalous window
# Input: daily cohort retention logs
# Output: which segments dropped, and by how much (table for LLM input)
def extract_break(rows, kpi="d30_retention", baseline_weeks=4):
    base = mean([r[kpi] for r in rows if r.week < target_week][-baseline_weeks:])
    cur  = [r for r in rows if r.week == target_week]
    return [
        {"segment": s.name,
         "baseline": round(base_by_seg[s.name], 3),
         "current":  round(s.value, 3),
         "delta_pct": round((s.value/base_by_seg[s.name]-1)*100, 1),
         "n": s.sample_size}          # sample size — small n means low confidence; pass it along
        for s in cur
    ]

The table this script spits out is the first input for the AI. The key point is that the sample size (n) travels with it. If you want the AI not to mistake the wobble of a small segment for a cause, the warning has to be carried by the data, not by a human.

# retention_break_2026Q2W3.txt (extraction result, excerpt)
segment                          baseline  current  delta_pct       n
new (signed up <7d)                  0.41     0.31     -24.4%   8,200
returning (30d+ dormant)             0.28     0.27      -3.6%   1,100
paying                               0.62     0.60      -3.2%   2,400
non-paying                           0.34     0.25     -26.5%  14,900
new_hunting_ground_played            0.39     0.22     -43.6%   3,050
new_hunting_ground_not_played        0.40     0.38      -5.0%  11,200

Step 2 — Prompt: Ban Diagnosis, Force Hypotheses and Verification Design

The attached retention_break_2026Q2W3.txt is the per-segment change in d30 retention
that Python extracted (baseline = average of the previous 4 weeks, current = this week,
n = sample size). Non-paying and new_hunting_ground_played dropped hard. Do not diagnose
the cause — build 3~5 verifiable hypotheses only. For each hypothesis, exactly these
four lines: the hypothesis in one sentence / the segment to verify (how to split it
further) / the pattern we'd see if it's right / the falsification condition that would
show it's wrong. Don't use small-sample segments (n<2000) as core evidence; if you do,
state the limitation. Sort the hypotheses by cheapest to falsify first, and mark
anything the data can't settle as 'needs human judgment'. Output only the list of
hypothesis cards, no narration.

There is not one line in this prompt that says "find the cause." Instead it specifies only the skeleton a hypothesis must have and the lines not to cross. Give the AI free rein and it will produce a plausible definitive sentence like "it's probably the new hunting ground's difficulty," and once that sentence lands in the meeting notes, it becomes a decision without verification. The guards against that are these four: no definitive statements; a mandatory verification segment, expected pattern, and falsification condition per hypothesis; sorting by falsification cost; and "what can't be settled goes to humans." They are what turn the output from a diagnosis into a verification design.

Step 3 — Claude's Output (Raw, Reconstructed)

[Hypothesis 1] (cheapest to falsify)
Hypothesis: the new hunting ground's difficulty became an entry barrier for non-paying new users and amplified churn within 7 days.
Verification segment: re-split the new_hunting_ground_played segment into 'early-death count' brackets (0~2 / 3~5 / 6+).
Expected pattern: if the hypothesis is right, the d7→d30 churn rate should spike in the 6+ early-deaths bracket.
Falsification condition: reject if there is no monotonic relationship between death count and churn rate.

[Hypothesis 2]
Hypothesis: the cause is not the new hunting ground itself but other changes bundled into the same patch (reward curve, UI).
Verification segment: retention of users who did not play the new hunting ground but logged in after the patch.
Expected pattern: if the not-played segment also fell, the hunting ground is not the culprit.
Falsification condition: if the not-played segment holds its baseline (around -5%), reject this hypothesis and move toward hypothesis 1.

[Hypothesis 3]
Hypothesis: an external factor (a competitor's new season) churned non-paying users across the board; the hunting ground is unrelated.
Verification segment: 2×2 cross table of paying/non-paying × played/not-played the new hunting ground.
Expected pattern: if it's external, non-paying users fall evenly regardless of whether they played the hunting ground.
Falsification condition: if the drop is concentrated in new_hunting_ground_played, the external-factor theory weakens.

[Hypothesis 4] (needs human judgment)
Hypothesis: a measurement problem — this week's cohort definition or log collection changed.
Verification segment: (data can't settle this) a human must check whether a telemetry schema change shipped with the patch.
Expected pattern: if the schema changed, the drop may be a measurement artifact rather than real.
Falsification condition: reject once the log pipeline is confirmed unchanged.

Note: the 'returning' segment (n=1,100) has a small sample, so its -3.6% was not treated as a signal.

The most valuable part of this output is Hypothesis 4 — the place where the AI reported its own limits: "the data can't settle this; a human needs to check whether the telemetry schema changed." A good prompt makes it possible for the AI to say "I don't know." And the last line — excluding the segment with a sample of 1,100 from the signals on its own — is the trace of rule 3 at work.

Step 4 — Verification and Decision (Where Humans and Python Belong)

You must not carry these hypothesis cards straight into a decision. A human sets the priority, and Python aggregates again.

Hypothesis 2 was the cheapest to falsify. The new_hunting_ground_not_played segment was already in the step 1 table — -5.0%. It held its baseline. In other words, users who skipped the hunting ground were fine. Hypothesis 2 was rejected on the spot, and hypothesis 3 (a broad drop from external factors) weakened at the same time. If the cause were external, the not-played segment should have fallen too. The drop was concentrated in users who played the new hunting ground.

So we narrowed to hypothesis 1 and ran Python again. Re-splitting new_hunting_ground_played by early-death count showed d30 churn standing out in the 6+ deaths bracket (direction: the more deaths, the steeper the churn — a monotonic relationship; exact figures measured by build telemetry, only the direction here). It matched hypothesis 1's expected pattern.

Hypothesis 4 remained. A human checked the patch notes — no telemetry schema change. The measurement-artifact possibility was rejected. Now the ingredients for a decision were in place.

[Step 5, Human Decision — Decision Card]

One cycle of input (anomaly signal) → extraction → hypotheses → verification → decision closes here. The AI never once said "the cause is this." It only laid the road for verification. This is the Show standard of this chapter — the sentence "AI analyzed the data" is empty unless you have watched, at least once and end to end, what was hypothesized, what was falsified, and what a human decided.


13.3.4 Why Definitive Diagnosis Is Banned

The difference between hypothesis generation and definitive diagnosis looks trivial, but it separates safe decisions from unsafe ones. Put side by side, the difference is plain.

Definitive diagnosis (banned) Hypothesis generation (this chapter's way)
AI output "The retention drop is caused by the new hunting ground's difficulty" "Difficulty hypothesis — look at the early-death 6+ bracket; if this, it's right; if that, it's wrong"
Human's next move Take dictation and decide Try to falsify, cheapest hypothesis first
When it's wrong The wrong decision ships straight into the build Rejected at the verification stage, cost 0
Accountability "The AI said so" (responsibility evaporates) A human picked the hypothesis and decided (responsibility is clear)

The real danger of definitive diagnosis is not its accuracy but that it makes people skip verification. One plausible sentence puts the room's doubts to sleep. A hypothesis card, by contrast, is itself homework — "go check this" — so structurally it cannot pass into a decision without verification. That is why I keep AI as a hypothesis generator, not a diagnostic machine.


13.3.5 Goodhart Early Warning — AI Flags KPI Distortion First

The deepest trap in data-driven work is Goodhart's law: "when a measure becomes a target, it ceases to be a good measure." Set DAU as the target and DAU alone inflates through artificial notifications while long-term retention erodes. The problem is that this distortion usually surfaces as side effects long after the decision was made.

So I bring AI in one slot earlier. Before a decision proposal goes into the build, I first ask the AI: "if we target this KPI, how could it be gamed?" This is not diagnosis — it's a red team. We deliberately make it hunt for the holes in our own decision.

[Goodhart early-warning prompt]

This quarter's target KPI is d7 retention +5%p, and the draft lever is a major buff to the 7-day consecutive login rewards. Act as the red team for this decision: give me a table of 3 Goodhart distortion scenarios that could arise from targeting this KPI, the guard metrics that would break alongside each scenario, and the monitoring segment that would catch the distortion early. No definitive claims — phrase each as "this could happen."

What the AI returned was not a definitive prophecy but a list of places to be suspicious of. The essentials:

Goodhart distortion scenario (hypothesis) Guard metrics that break with it Early monitoring
Logging in for the stamp without playing core content Battles per session, hunting ground entry rate Alert when d7 retention ↑ and battle count ↓ occur together
Reward inflation breaking the economy Currency sink/source ratio, item market prices Track the widening sink–source gap in _economy_log
Cliff churn right after the login streak ends d8–d14 retention (right after rewards end) Don't watch d7 alone — pair it with d14

The value of this table is not that it is correct but that it pairs up guard metrics before the decision. If we're going to target d7 retention, we put the "battle count" and "d14 retention" the AI flagged on the same screen and watch them together. Then the moment d7 rises while battle count falls — the moment Goodhart distortion begins — gets caught before the side effects accumulate to quarter's end. This habit of pairing a KPI with guard metrics, instead of targeting a single KPI, is how the "5–7 KPI balance" set in 13.2 actually operates at the decision stage.

One thing worth pinning here. The value the AI created in this red team is not "time saved." It does not take a human long to think of these three scenarios. The real value is that it exposes the distortion signals right where the decision is being made — the signaling effect of pulling guard metrics nobody usually watches onto the decision table. The value of automation lies not in saving time but in making normally invisible signals visible (Project A team memory concept automation_signal_value_over_time_savings).


13.3.6 AI Hypotheses Carry Different Weight for Different Decisions

Hypothesis generation is not equally useful for every decision. How much to trust AI hypotheses changes with the decision's time horizon and data density.

Decision type Data density Where AI hypotheses stand
Skill balance number change High (rich sims and logs) Run the hypothesis→verify→decide loop as is; AI assist is strong
UI component change High (A/B testable) Same; AI hypotheses valid
Whether to ship new content Medium (only similar-content references) Hypotheses are reference material; decision weight shifts to humans
Long-term vision, new domains Low (no precedent) The loop itself doesn't run — humans decide; AI only enumerates risks

The rule is simple. The thicker the data behind a decision, the more you run the §13.3.2 loop as is; the thinner the data, the more AI steps down from hypothesis generator to risk-checklist writer. Trying to solve long-term vision with data is dangerous because, where future data does not exist, an AI will fabricate plausible hypotheses out of past data, and those hypotheses drag the vision back toward the past. Decisions in data-free territory are not to be dodged or offloaded onto AI — they remain the seat where a human takes responsibility and decides.

[Signpost — If Embeddings Could Map Topics and Cohorts into Coordinates (Still Premature)]

Read this as a research trend, not a prescription. The same embedding idea opens up in two places in Part 13. One is the free-form responses of §13.1 — clustering unstructured natural language with sentence embeddings would let you express the [ambiguous] boundary cases of §13.1.2 as "distance between two topic centroids," and flag responses far from every centroid as "a new topic emerging." The other is the behavior logs of §13.1.4 — embedding play logs could reveal "emergent cohorts" nobody predefined, as clusters in vector space (the "map" of Appendix M), feeding them in as candidate "segments to verify" for the §13.3 hypothesis loop (it punches one hole in the limitation §13.3.3 assumed: segments predefined by humans). But a cluster is a hypothesis, not a cause; small clusters are not signals (the same seat as §13.3.3's sample-size warning); and naming the clusters — the labeling — is still human work (§13.1.1). Above all, a live incident can erupt in a dimension the compression threw away. So I file this idea in exactly the same place as the "dimension vector" lead in the economy chapter, §8.2.7 (conceptual intuition in Appendix M) — on the same telemetry soil, with the same restraint. It is a signpost for teams with telemetry solidly laid down to revisit a few years from now; the thing to do today is to run the §13.3.2 loop honestly.


13.3.7 Where This Chapter's Numbers Come From

The numbers in this chapter follow the principle of the preface's "One Promise." Goodhart's law is a public proposition formalized by Charles Goodhart in 1975; Project A's _economy_log, _roi_report.md, and _scores_latest.json are real team memory artifacts; and the rule that notifies via ClickUp on an integrity failure, integrity_check_clickup_notify, is a live operational atom with a score of 294.93 (Appendix A.3.6, A.3.1). In §13.3.3, only the direction — "churn is steeper in the early-death 6+ bracket" — was confirmed through hypothesis verification; absolute values were left to build telemetry. The segment table (baseline 0.41 and so on) is an illustrative construction to show the shape of the workflow, not published measurements from a specific quarter — what to memorize is the structure, not the numbers.


13.3.8 Common Failures

Pattern Why it fails Remedy
Asking the AI "what's the cause" A plausible definitive sentence becomes a decision without verification Ban diagnosis; force 3–5 hypotheses + falsification conditions (§13.3.3)
Treating a small segment's wobble as a signal Mistakes noise for a cause Pass n along at the extraction stage and state the threshold
Going straight to a single-KPI target Goodhart distortion erupts at quarter's end AI red team before deciding + paired guard metrics (§13.3.5)
Deciding long-term, data-free questions with data Past-built hypotheses drag down the future vision Tier the AI's role by data density (§13.3.6)
Accepting hypotheses without verification A hypothesis masquerades as a conclusion Falsify cheapest-first; use the not-played segment

The third one detonates last. d7 retention rises, the decision looks like a success, and two months later the d14 cliff and the battle-count drop arrive together. The 30 minutes spent running the AI red team before the decision buys back those two months.


13.3.9 Try It Yourself — One Step You Can Take Today

If you're solo, just this much: you don't need a log pipeline. Pick one recently bent number from your own game (or from the public metrics of a game you follow). Toss that number to the AI — but instead of "tell me the cause," ask for "no definitive diagnosis; 3 verifiable hypotheses, each with a falsification condition." Pick the one hypothesis that is cheapest to check and split the data yourself, once. You will feel in your bones how different "receiving a diagnosis" and "verifying a hypothesis" are for the safety of a decision.

If you're on a team, start with this one step. When your anomaly extraction script emits per-segment numbers, add one line so it always outputs the sample size (n) too (the retention_break_extract.py of §13.3.3). And the next time you set a KPI target, run the Goodhart red-team prompt of §13.3.5 once and enter one pair of guard metrics into the decision card. With just these two, "the AI diagnosed the cause" turns into "the AI laid out hypotheses, and a human verified and decided."


Key Takeaways

Next Chapter Preview

Part 14 · Mobile Platform

14.1 From 30 PC HUD Elements to 10 on Mobile — Constraints as a Rulebook, Compression by AI

Primary audience: UX and systems designers on mobile-first projects (mid-size teams of 10–50) Scaled-down version for solo/hobbyist readers: §14.1.7, "If You're Solo, Just This Much"

I remember the day we first put the combat HUD that ran fine in the PC build onto a mobile screen. Half the screen was buried under gauges, icons, the minimap, and the quest tracker — and the character itself was nowhere to be seen. Every single element looked necessary. The problem was that "what do we cut" turned back into a from-scratch fight at every meeting. Someone wanted to save the minimap; someone else wanted to save the chat window. Because the rationale was "feel," the conclusion came out different every time.

This chapter is about ending that fight. There are two keys. First, turn mobile constraints from "feel" into a verifiable rulebook. Second, hand the tedious, repetitive compression work — cutting 30 PC elements down to 10 for mobile — to the AI, and have humans do only the review that catches rulebook violations. General mobile UX knowledge already fills plenty of other books, so this chapter focuses solely on running that knowledge through an AI workflow.


14.1.1 Mobile Constraints Are a Rulebook, Not a List of Cautions

Plenty of books lay out mobile constraints in a table. The screen is small, fingers are thick, sessions are short, the battery drains. All true — but memorizing that table won't answer the meeting-room question, "So is this button OK or not?" Constraints have to become numeric pass/fail criteria before AI and humans draw the same line.

Fortunately, most mobile input constraints have already been nailed down by platform vendors as public guidelines. Public standards like the 44pt touch target (HIG), 48dp (Material), 4.5:1 contrast (WCAG), and 8dp spacing follow the §9.1 rulebook; here I keep inline only the one this chapter's lint uses directly — the 44pt minimum touch target (HIG). These are numbers nobody needs to make up. Only when you can say "this button is 38pt, below the HIG 44pt" instead of "this button feels a bit small" do you get the same verdict whether a human or an AI is making the call.

I add one more line on top — for mobile MMORPGs, the landscape two-handed grip is the standard: pressable elements go in the two bottom corners, and consumables/slots go in the bottom center (why landscape is the standard and what the three-zone model is are covered in §9.1). Every placement verdict in this chapter assumes that landscape two-handed grip.

Put the platform criteria side by side with PC, and the starting point of compression becomes obvious. PC is precise and high-capacity (it can carry 30–50 elements); mobile landscape is confined to the two-thumb corners, so 12–16 is the ceiling (full comparison table in the §9.1 rulebook — author's estimate, unverified). So the essence of mobile work is not "design" but "priority-compressing 30–50 PC elements into the 12–16 of mobile landscape." And this compression is tedious by hand, and the baseline drifts every time you redo it — it is the job of applying the same rules over and over without tiring, which fits exactly the division of labor where AI drafts and a human reviews.


14.1.2 [Worked Transcript] 30 PC HUD Elements → Mobile Priority Compression

Here is one full cycle of how I actually run it. Below is a faithful reproduction of a combat-HUD compression session from my own project (a mobile-first MMORPG, "Project A" hereafter). The input prompts can be copied and used as-is; the outputs are reconstructions of the actual session.

Step 1 — Input: Feed the PC HUD Spec as Is

First, turn the PC HUD element list into a machine-readable table. This already lives in the data sheets, so it is extraction, not new writing.

# hud_pc_inventory.yaml — current HUD in the PC build (excerpt, 12 of 30 elements)
- id: hp_bar          # HP bar
  현재위치: top-left
  상시노출: true
  조작가능: false
- id: mp_bar          # MP bar
  현재위치: top-left
  상시노출: true
  조작가능: false
- id: skill_slots     # 12 skill slots
  현재위치: bottom-center
  상시노출: true
  조작가능: true
- id: minimap         # minimap
  현재위치: top-right
  상시노출: true
  조작가능: true
- id: quest_tracker   # quest tracker
  현재위치: right side
  상시노출: true
  조작가능: false
- id: chat            # chat window
  현재위치: bottom-left
  상시노출: true
  조작가능: true
# ... buff_bar, party_frame, target_frame, exp_bar, currency, mail_alert ...

Step 2 — Prompt: Nail Down the Classification Format and the One-Line Rationale

Take the attached hud_pc_inventory.yaml (the 30 current HUD elements in the PC build)
and priority-compress it for mobile landscape two-handed grip. Classify each element
from P0 (always required during combat) to P3 (remove or contextual exposure), and
keep the always-on total (P0–P1) at 16 or fewer. For interactive elements
(조작가능:true), propose placement in the thumb-reach corners (bottom-left = movement,
bottom-right = skills). Attach a one-line rationale to every classification, and don't
force a call on ambiguous ones — flag them separately and hand them to a human.
Output the table only (element | class | zone | rationale | ambiguous), no commentary.

Note that in this prompt, the instruction to compress is a single line; everything else is fencing that pens in the output. Give the AI freedom and it regurgitates the generic-RPG average. So I pre-load the classification criterion (P0–P1 ≤ 16), the output format (a table), the rationale requirement, and the escape hatch — "hand ambiguous ones to a human." The quality of the compression is decided by these four fences.

Step 3 — Claude's Output (Reconstruction, Excerpt)

| Element     | Class | Zone                    | Rationale                                    | Ambiguous |
|-------------|-------|-------------------------|----------------------------------------------|-----------|
| hp_bar      | P0    | top, fixed              | death-critical, needs checking every 0.5s    |           |
| skill_slots | P0    | bottom-right (R thumb)  | core controls, must sit in right-thumb corner|           |
| mp_bar      | P0    | top, fixed              | decides whether skills can fire, always needed|          |
| target_frame| P0    | top center              | target HP check, core combat read            |           |
| minimap     | P1    | top-right, shrunk       | direction check, can shrink/fade in combat   |           |
| buff_bar    | P1    | below hp_bar            | buff-expiry check, icons can shrink          |           |
| quest_tracker| P2   | menu toggle             | not needed in combat, toggle outside combat  |           |
| chat        | P2    | bottom-left tab         | auto-hide in combat, tap to expand           | ★         |
| party_frame | P1    | left side, shrunk       | promote to P0 in party play (conditional)    | ★         |
| mail_alert  | P3    | contextual              | shows briefly only on receipt                |           |
| exp_bar     | P3    | bottom 1px line         | low info value, a thin line is enough        |           |
| currency    | P2    | inside menu             | not needed in combat                         |           |

Always-on total (P0+P1): 7 (rule 2 pass: ≤16)
2 ambiguous items (chat, party_frame) need a human decision:
- chat: consider promoting to P1 if the game leans heavily social
- party_frame: classification splits by solo/party context

The most valuable part of this output is not the table but the two lines at the bottom — "2 ambiguous items." That is the AI reporting its own limits and handing the judgment to a human. Leave an escape hatch in the prompt, and instead of fabricating an answer, the AI plants a flag where it doesn't know.

Step 4 — Verification and Veto (the Human's Job)

Don't accept this output as-is. Run it against the rulebook once. In this actual session, one item got caught.

The AI placed party_frame at "left side, shrunk" — but in a landscape grip, the left-center of the screen is a zone neither thumb reaches (the left hand is tied to bottom-left movement, the right hand to bottom-right skills). Yet the party frame is an interactive element that needs taps (targeting party members). That violates rule 3 ("interactive elements go in the easy thumb corners"). The AI missed the 조작가능 flag on party_frame. The cause: party_frame's 조작가능 field was empty in the input yaml — a data defect on the human side.

So I re-request.

party_frame is an interactive element that needs taps for targeting party members
(it was missing from the earlier input). Re-place it under the rule that interactive
elements must sit in the thumb-reach corners. Propose separately for solo play and
party play.

One round trip and it's done. The AI answered again with "hidden" for solo play and "promoted to bottom-right (easy)" for party play, and that decision passed the rulebook. Compressing 30 elements by hand from scratch takes half a day; AI draft + rulebook review + one round trip stays under an hour (author's estimate — the exact time saved varies by team and element count, so read it less as an absolute value and more as the structural difference between "by hand from scratch" and "draft + review").


14.1.3 Finger Zones — Both Corners and the Bottom Center

Pin down the "finger zones" that kept coming up in the session above as one diagram, and every placement verdict afterward gets faster. On a phone held in landscape, the bottom — where fingers reach and eyes go often — splits into three spots. The left thumb reaches the bottom-left corner (movement), the right thumb the bottom-right corner (skills), and the bottom center between the two thumbs is where consumables, auto-use items, and skill slots go. It isn't twitch input, but it is an important glance zone where you see at a glance what you use or what auto-consumes, and press it occasionally. P0 controls and slots are green; the top and upper center — where fingers never go and you only read — are red.

Hard — top & upper center (status only: HP · MP · target, read-only) Game view (where the combat happens) L thumb Move R thumb Skills Bottom center — consumables · quick slots · auto Potion Auto Slot HP MP TGT Map Move Skill Skill Skill

The rule is simple. Read-only information (HP/MP/target health) can live in the red (top and upper center) — fingers never need to go there. Conversely, pressable elements must be inside the finger zones (green and amber) — movement and skills in the two bottom corners; consumables, auto-use items, quick slots, and skill slots in the bottom center. All three are spots where fingers reach and eyes visit often. This one picture explains why party_frame got caught in §14.1.2 — a pressable element was placed in the left-center (a reading zone), not in a finger zone.


14.1.4 The Rulebook as Code — Automated Layout Lint

Check by eye every time whether a compressed plan honors the rulebook, and you will miss things again. Of the five rules in §14.1.1, the ones decidable by coordinates and size should be reviewed by code. Humans spend time only on the "ambiguous" calls that code can't make.

# hud_lint.py — mobile HUD layout-plan verification (skeleton)
# Input: the AI-proposed layout plan (per-element coordinates, size, 조작가능, 분류)
# Output: list of rulebook violations

MIN_TAP_PT = 44       # Apple HIG minimum touch target (pt)

def in_action_zone(e, w, h):
    """Finger-reach zones in landscape grip: bottom-left/right corners + bottom-center slot bar."""
    x, y = e["x"] / w, e["y"] / h
    bottom = y > 0.55
    left_corner  = bottom and x < 0.30                 # left thumb = movement
    right_corner = bottom and x > 0.70                 # right thumb = skills
    center_slot  = (y > 0.72) and (0.35 <= x <= 0.65)  # bottom center = consumables/quick slots
    return left_corner or right_corner or center_slot

def lint(elements, screen_w, screen_h):
    issues = []
    for e in elements:
        # Rule A: interactive/slot elements must sit in the finger zones (both corners + bottom center)
        if e["조작가능"] and not in_action_zone(e, screen_w, screen_h):
            issues.append(f"[A] {e['id']}: interactive/slot element placed outside finger zones "
                          f"(x={e['x']}, y={e['y']})")
        # Rule B: minimum touch-target size (HIG 44pt)
        if e["조작가능"] and min(e["w"], e["h"]) < MIN_TAP_PT:
            issues.append(f"[B] {e['id']}: touch target {min(e['w'], e['h'])}pt "
                          f"< {MIN_TAP_PT}pt (below HIG)")
    # Rule C: total always-on P0/P1
    onscreen = [e for e in elements if e["분류"] in ("P0", "P1")]
    if len(onscreen) > 16:
        issues.append(f"[C] {len(onscreen)} always-on elements > 16 (overcrowded)")
    return issues

With these 30 lines, "isn't this button a bit small?" stops being a discussion topic in meetings and becomes a verdict. When the code prints [B] skill_slots: touch target 40pt < 44pt (below HIG), there is no need to gather opinions. You fix it. This is the lint gate from 9.1 (HUD) carried over to the mobile dimension — the division where code catches what is deterministically decidable and humans take what is non-deterministic and judgment-based holds on mobile just the same.

Here is the full cycle at a glance.

flowchart LR
    A["30 PC HUD elements
(extracted from data sheets)"] --> B["AI compression
P0–P3 classification + placement"] B --> C{"hud_lint.py
automated rulebook check"} C -->|violation| D["Re-request
(fix omissions, misplacements)"] D --> B C -->|pass| E["Human review
'ambiguous' calls only"] E --> F["Mobile HUD finalized
around 12–16 elements"] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; class C code; class B ai; class E human; class A data; class F pass;

Human hands touch only two places: putting clean input data in (the very front), and making the ambiguous judgments the rulebook can't catch (the very back). The tedious 30-element compression in between is run by the AI and the lint.


14.1.5 Where the Numbers in This Chapter Come From

A short record of the sources for the numbers in this chapter (for the book-wide policy on numbers, see "One Promise" in the preface). The 44pt touch target (HIG), 48dp (Material), and 4.5:1 contrast (WCAG) are official platform standards; "8–12 always-on info elements" and "compression: half a day → one hour" are the author's experience-based estimates (unverified), so read them as direction rather than absolute values. The metrics actually measurable on a mobile HUD are rulebook violation count (lint 0), always-on element count (target ≤12), and mis-tap rate (telemetry); outcome metrics like retention are not decided by the HUD alone, so I don't assert causation.


14.1.6 Common Failures

Pattern Why it fails Fix
Porting the PC HUD shrunk as-is 30 elements bury a 6-inch screen; the game disappears The compression session in §14.1.2
Wholesale delegation — "AI, build me a mobile UI" Without a rulebook you get the generic-RPG average Feed the rulebook (§14.1.1) into the prompt first
Reviewing the compressed plan by eye only Touch-size and thumb-zone violations slip through every time Automated verification with hud_lint.py
"Let's just cut this" meetings with no rationale The conclusion changes every time Force P0–P3 + a one-line rationale

14.1.7 Try It Yourself — One Step You Can Take Today

If you're solo, just this much: You don't need data sheets. Write down 10–15 PC HUD elements of your own game (or a game you love) by hand, turn them into yaml, paste in the prompt from §14.1.2 as-is, and run it once. Find one item where you disagree with the AI's classification and push back — "justify this again" — and you'll feel in your bones what bundle of judgments compression really is.

If you're on a team, start with this one step. Extract the current HUD element list into hud_pc_inventory.yaml (it's already in your data sheets), and pin down the three rulebook lines of §14.1.4's hud_lint.py (touch size, thumb zones, total count) as code first. With the rulebook in place, you can measure an AI compression plan and a human draft against the same line.


Key Takeaways

Next Chapter Preview

14.2 Platform Differences (iOS / Android / PC)

The day we first put the alpha build on PC, a screenshot appeared in the design team's messenger channel. The virtual joystick that filled the bottom of the screen on mobile was floating, palm-sized, in the middle of a 27-inch monitor. Someone added one line: "How am I supposed to grab this with a mouse?" The core logic was fine. Combat, inventory, quests — all of it ran as before. Exactly one thing had broken: the places where input and screen had been pinned to mobile assumptions.

Shipping the same game to iOS, Android, and PC looks like it should triple the operational load, but in practice it does not. There is one core logic, and three platform adaptation layers attach to it. The problem is that "where the core ends and the adaptation layer begins" is hard for a person to judge case by case. A branch that works on iOS but breaks only on Android, a key mapping that only matters on PC — these differences do not all fit in one head. So the heart of this chapter is a workflow that codifies platform constraints into a rulebook, has AI generate branch drafts grounded in that rulebook, and finally has lint catch the rule violations.


14.2.1 What Differs Across the Three Platforms

First, the terrain of the differences. Below is the platform constraint table we put together on Project A (a mobile-first MMORPG where I serve as design director) while evaluating a secondary PC release. Where a number is based on a public standard, I noted the source; the rest are values agreed internally on the project.

Area iOS Android PC
Input Touch Touch (+ some keyboards) Keyboard, mouse, gamepad
Minimum touch target 44pt (Apple HIG) 48dp (Material) Click — not applicable
Screen 4.7–6.7 inches 4.5–7 inches (high variance) 21–32 inches
Payments App Store Google Play In-house / Steam
Notifications APNs FCM OS / in-house
Save data iCloud Google Drive / in-house Steam Cloud / in-house
OS replacement cycle 1–2 years 1 year (heavy fragmentation) 5–10 years

iOS and Android differ in their payment, save, and notification APIs, but what the user sees and how they control the game are almost identical. PC differs wholesale in input, screen, and visual effects. So the operational load, counter to intuition, is closer to ×2 than ×3 — because the distance between iOS and Android is short.

What matters here is not the table itself but turning it into a rulebook that machines read, not a document that humans read. That is what lets AI use it as grounding when generating branch drafts, and lets lint catch violations.


14.2.2 The Line Between Core and the Platform Layers

Project A's folder structure attaches three platform adaptation layers to a single core.

game/
├── core/                  — game logic (platform-agnostic)
│   ├── combat/  inventory/  narrative/  ...
├── platform/              — platform adaptation layers
│   ├── ios/      → input/  payment/  notification/
│   ├── android/  → input/  payment/  notification/
│   └── pc/       → input/  payment/  ui/
└── shared/                — used by all (utilities · rendering)

There is one rule. core never calls platform by name. The moment core contains a statement like if platform == "ios", the layer separation collapses. Take input: core only knows the intent "use skill 1" (InputIntent.SKILL_1); whether that intent is extracted from touch coordinates or from the keyboard's 1 is the responsibility of each platform layer.

Draw this line and the next step becomes possible: adding a new platform means filling in a single folder under platform/ without touching core. The figure below shows, on one page, how this line actually splits.

core/ Game logic · platform-agnostic InputIntent · PaymentInterface platform/ios touch → intent StoreKit · APNs targets ≥ 44pt iCloud saves platform/android touch → intent Play Billing · FCM targets ≥ 48dp fragmentation handling platform/pc key/mouse → intent Steam · OS notifications gamepad · key-mapping UI varied resolutions shared/ — utilities · rendering PC differs wholesale in input, screen, and visuals (orange)

The iOS and Android boxes share the same blue family and only PC is orange — the size of the difference is shown in color. The asymmetry of the operational load is visible here at a glance.


14.2.3 The Rulebook: Making the Differences Machine-Readable

The key turning point is here. Write platform constraints into a prose document and people forget them. Instead, gather them into a single declarative rulebook file. Below is an excerpt from the platform_rules.yaml we use on Project A (I trimmed the actual file down to the core rules for this chapter).

# platform/platform_rules.yaml
targets:
  ios:
    min_touch_pt: 44          # Apple HIG
    contrast_ratio: 4.5       # WCAG SC1.4.3
    gamepad: optional         # iOS 17+ standard
    forbidden_in_core: ["import platform.ios", "StoreKit", "APNs"]
  android:
    min_touch_dp: 48          # Material
    contrast_ratio: 4.5
    forbidden_in_core: ["import platform.android", "BillingClient", "FCM"]
  pc:
    min_target_px: 24         # WCAG SC2.5.8 (pointer)
    input: ["keyboard", "mouse", "gamepad"]
    forbidden_in_core: ["import platform.pc", "SteamAPI"]
required_intents: ["MOVE_FORWARD", "ATTACK", "SKILL_1", "SKILL_2"]

This one file does three things at once: (1) the spec AI reads when generating branch drafts, (2) the criteria lint verifies against, and (3) the single source where people record what was agreed. forbidden_in_core matters most — it is the list of tokens that must never appear inside the core folder, which gives lint a mechanical basis for catching layer violations.


14.2.4 Worked Transcript: Rulebook → AI Branch Draft → Verification

Now I walk through a real task from start to finish. The situation: preparing the secondary PC release, an input handler that existed only on mobile needs a PC branch. Grounded in the rulebook, I get a draft PC input adapter from Claude, verify it with lint, and complete one full cycle including rejection and re-request.

1) The Prompt (Verbatim)

Below are an excerpt from our project's platform rulebook and the existing iOS input handler.

[the platform_rules.yaml excerpt above, pasted in]

# platform/ios/input/touch_handler.py
from core.input.input_intent import InputIntent

def map_touch_to_intent(touch_event):
    if touch_event.location in skill_button_1_area:
        return InputIntent.SKILL_1
    if touch_event.is_drag_up:
        return InputIntent.MOVE_FORWARD
    return None

Using this as reference, draft platform/pc/input/keyboard_handler.py. Constraints: (1) import core, but never import platform.ios/android, (2) map all 4 required_intents, (3) separate the key mapping into a dict so users can change it. Flag anything you are not sure about.

2) Claude's Raw Output (No Summarizing, No Cleanup)

Here is a draft. I pulled the key mapping out into a dict at the top of the module so users can customize it.

# platform/pc/input/keyboard_handler.py
from core.input.input_intent import InputIntent

DEFAULT_KEYMAP = {
    "w": InputIntent.MOVE_FORWARD,
    "1": InputIntent.SKILL_1,
    "2": InputIntent.SKILL_2,
    # ATTACK assumed on left mouse click — no keyboard default assigned
}

def map_keyboard_to_intent(key_event, keymap=DEFAULT_KEYMAP):
    return keymap.get(key_event.key.lower())

def map_mouse_to_intent(mouse_event):
    if mouse_event.button == "left":
        return InputIntent.ATTACK
    return None

⚠️ Where I am not confident: required_intents includes ATTACK, and I mapped it to the mouse, not the keyboard. The rulebook does not specify whether required_intents means "OK as long as it is mapped on some input device" or "everything handled within a single handler," so I assumed the former. If it is the latter, lint may flag ATTACK as missing.

3) Human Verification / Rejection

The ⚠️ spot Claude flagged was exactly the trap. Our lint (platform_lint.py) checks required_intents coverage per platform, not per handler file. Even if ATTACK is missing from keyboard_handler.py, it passes as long as it exists on the mouse_handler side. But the output Claude produced put the mouse mapping inside the keyboard_handler.py file — file responsibilities got mixed. Structurally it would pass, but it violates our folder rule (separate files per input device). Rejected.

The rejection reasons fit in two lines: (1) split the mouse mapping into a separate mouse_handler.py. (2) Make ATTACK usable from the keyboard too, with Space as a fallback.

4) Re-request → Lint Pass

After the re-request, I ran the split version through platform_lint.py. Lint reads the rulebook and checks the following.

$ python platform_lint.py platform/pc/
[core-leak]    PASS  — 0 forbidden tokens inside core/
[intent-cover] PASS  — pc: MOVE_FORWARD, ATTACK, SKILL_1, SKILL_2 (4/4)
[touch-target] SKIP  — pc uses min_target_px=24 (checked separately in the UI layer)
[no-cross-import] PASS — platform.pc does not reference platform.ios/android

intent-cover landing at 4/4 is the point. Whether the AI-generated draft meets the rulebook's criteria was settled by a script, not by human eyes. That one line replaces the re-checking a person used to do in their head, every time, in multi-platform operations.

Compressed into a diagram, the cycle looks like this.

flowchart LR
    R[platform_rules.yaml
rulebook] --> P[Prompt with
rulebook + existing handler] P --> A[Claude branch draft
+ uncertainty flags] A --> H{Human verification} H -->|Reject: mixed file responsibilities| P H -->|Accept| L[platform_lint.py] L -->|FAIL| P L -->|PASS| M[To build branching] R -.provides criteria.-> L classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; class L code; class A ai; class H human; class R data; class M pass;

The center of this structure is that the rulebook supplies the criteria to both the prompt and lint. AI generates, a human judges, lint confirms — all three roles look at the same rulebook.


14.2.5 Build Branching: Same Core, Different Assembly

Once the handlers are in place, the build is simple assembly. core and shared are fixed; only the platform folder swaps out.

[core/ + shared/ + platform/ios/]      → iOS build
[core/ + shared/ + platform/android/]  → Android build
[core/ + shared/ + platform/pc/]       → PC build

In CI we run these three in parallel, not sequentially, and run platform_lint.py automatically right after each build. Run them sequentially and build time triples; drop the lint and rule violations survive all the way to deployment. Parallel builds plus automatic lint — those two are the minimum requirements of multi-platform CI.

Release cycles differ per platform, so a passing build does not mean simultaneous deployment. iOS review usually takes 1–3 days, which makes frequent releases a conservative bet; Android propagates within hours, so you can ship more often; Steam sits around 1–2 days. The same change ships last on iOS, so hotfix schedules are always back-calculated from iOS.


14.2.6 UI Variants: 80 Common · 15 Variant · 5 Exclusive

Beneath the code, the screens split too. In my experience, the recommended distribution is 80% common components, 15% platform variants (differing only in size and position), and 5% platform-exclusive. This ratio shifts with genre, though — a casual puzzle game pushes common up to 90%, while an MMORPG grows more variants because of the input differences.

Exclusive components are where a platform's appeal lives, so forcing everything into common is not the answer. Mobile's virtual joystick and vibration, PC's key-mapping UI and gamepad settings — things meaningful only on that platform belong here. But once exclusive passes 30%, that is a signal of operational load, not appeal — put a platform-specific-ratio warning in lint and the build will point it out even when people forget.

This is also the boundary of AI assistance. Platform differences are mostly deterministic rule territory, so AI is used to generate branch drafts that satisfy the rulebook rather than to explore candidates freely. Input mapping recommendations, converting Figma mocks into platform variants, adapting text across languages × platforms — that is roughly where AI genuinely adds value, and its output must always pass lint. Adapter standardization comes before progressive automation.


14.2.7 What the Separation Is Worth — And the Common Pitfalls

The biggest payoff of layer separation is the speed of adding a new platform. Pile if-statements onto a single code base to bolt on PC and the cost approaches building a new game; fill in platform/pc/ without touching core and that time shrinks dramatically. The speedup ratio varies by project, so I will not assert a specific multiplier — but in our internal review, we estimated the secondary PC schedule at less than half of the single-codebase assumption (author's estimate, unverified). As side effects, platform-specific incidents stay isolated, and the reliability of core changes goes up (fix one place and it propagates consistently to all three builds).

The pitfalls we step on most often, and their prescriptions:

Pitfall Prescription
if platform == ... branches multiplying inside core Block with the forbidden_in_core lint; separate into adapters
Reviewing AI branch drafts with human eyes only Confirm intent-cover with platform_lint.py
Cramming all input device mappings into one file Separate handlers per device (keyboard/mouse)
Exclusive components at 30%+ platform-specific-ratio warning; consider commonizing
Deploying to all three platforms the moment the build passes Back-calculate from iOS, following the release-cycle differences

What the pitfalls share is that "someone tried to hold the line from memory." Write it in the rulebook and wire it into lint, and the build remembers even when people forget.


Key Takeaways

Next Chapter Preview


Try It Yourself

setup. Create platform/platform_rules.yaml in your project and, like the excerpt above, write per-platform min_touch, contrast_ratio, forbidden_in_core, and required_intents. Do not make the numbers up — take them from public standards (for public standards like the 44pt/48dp touch targets and the 4.5:1 contrast ratio, follow the §9.1 rulebook; the PC pointer target of 24px is WCAG SC2.5.8).

prompt. Paste the rulebook excerpt together with one existing platform's handler, and ask: "Follow this rulebook and draft a platform/<new-platform>/input/ handler. Never include forbidden_in_core tokens, map all of required_intents, and mark anything you are not sure about with ⚠️."

verify. Run a platform_lint.py (a 40-line script is enough) that reads the rulebook and checks: (1) 0 forbidden_in_core tokens inside the core folder, (2) all required_intents mapped per platform, (3) no cross-imports between platform folders. If even one check FAILs, go back to the prompt, write down the rejection reason, and re-request.

Solo Scale-Down

If you work alone and have no build CI, shrink the rulebook from YAML to a one-page markdown checklist. Three lines will do: "targets ≥ 44pt, no platform imports in core, all 4 intents mapped." Instead of a lint script, hand the AI your result and tell it: "Judge each of these 3 checklist items as pass or fail, one by one." That stands in for the human re-check. The point is not the scale of the tooling — it is writing the criteria down outside your head, and separating generation from verification.

14.3 Touch / Mouse Input Design

Team member B picked up the QA build, gripped the phone in one hand, and frowned. "I tapped the skill three times, but it only fired twice." Looking at the screen, at the exact moment the thumb pressed the skill button, that same finger was covering a third of the area next to it. The problem never once appeared when testing with a mouse. A mouse has no fingers.

That scene sums up the essence of touch versus mouse in one line. Both are inputs that "point at a spot," but one pointing tool covers the screen and the other does not. This is where the need to solve the same action differently for the two inputs begins. In this chapter I first sort out the differences between the two inputs, then follow a single worked-transcript spine all the way through: getting an input mapping proposed by AI, then verifying conflicts and reachability myself.


14.3.1 The Fundamental Differences Between the Two Inputs

Finger thickness, screen occlusion, multi-touch limits, and precision all differ. Before pinning this down in a table, let's get a feel for it with one image. A mouse cursor is a 1-pixel pen nib; a finger is a stamp nearly a centimeter across. The nib can write, but one letter at a time. The stamp prints fast, but it cannot write — and the moment it presses down, you can no longer see the paper.

Property Touch Mouse
Precision About 7–10mm (finger contact area) 1px granularity
Occlusion Finger covers the area around the contact point None
Hover Nearly impossible (contact = input) Free (movement ≠ input)
Simultaneous inputs 2–10-point multi-touch Left, right, middle, wheel
Drag/tap distinction Must be inferred from time and distance Click/drag unambiguous
Haptic feedback Available Mostly none

The two rows with the biggest design impact are "occlusion" and "hover." Occlusion dictates where results get displayed, and the absence of hover means an entire information channel — the tooltip — disappears on mobile. The other four rows are mostly details derived from these two.

Public standards nail these differences down in numbers — open standards like 44pt touch targets (HIG), 48dp (Material), 4.5:1 contrast, and 24 CSS pixels for touch targets (WCAG SC2.5.8) follow the §9.1 rulebook. These numbers are products of human anatomy and measurement, not taste, so when it comes time to verify a mapping, the yardstick you hold up is ultimately these same standards.

14.3.2 Mapping by Game Action

Movement, attack, and skills — resolving these three actions for the two inputs splits as follows. Having three options per action does not mean there is no right answer; it means the game's identity forces the choice.

On Project A, the mobile-first MMORPG I work on, movement uses an ⓐ+ⓑ hybrid on mobile (joystick and auto-move side by side), and WASD plus auto-move on PC. Attack is ⓑ+ⓐ on mobile (tap enemy, then button); PC offers a free choice of ⓐ or ⓑ. Skills are ⓐ on mobile, or ⓑ when targeting; PC layers mouse aiming on top of keys 1–8. Same game, same actions — and yet the mapping table comes out as two sheets. That fact is this entire chapter.

The problem is that the longer the mapping table grows, the better conflicts hide. What happens when long-pressing a slot (skill aiming) and dragging the screen (camera rotation) overlap in the same region? Keys 1–8 are taken by skills, and someone proposes 1 as the inventory shortcut too? Scanning the table by eye, you miss it. That is why we need a workflow where AI proposes the mapping but verification is handed to a tool.

14.3.3 Worked Transcript — AI Mapping Proposal and Lint Verification

From here on is the spine of this chapter. I carry over one full cycle in its entirety: getting a draft PC keyboard/mouse mapping proposed by Claude, having a human verify the output, running a conflict-and-reachability lint, rejecting, and re-requesting. I do not summarize. Raw output contains mistakes, and the point is to watch the process of catching them.

The Prompt (in Full)

You are designing the input mapping for a PC MMORPG. Propose JSON that maps the action list below to keyboard/mouse.

Actions: movement (forward/back/left/right), run, jump, basic attack, skills 1–8, inventory, map, character window, menu, camera lock toggle, zoom in/out, context menu (secondary action), auto-move toggle.

Constraints: - Follow industry conventions (WASD movement, ESC menu, etc.). - Do not assign two actions to a single physical key/button. - Each entry has action, input, note fields. Output the JSON array only.

Claude's Raw Output (Verbatim, Untouched)

[
  {"action": "move_forward",  "input": "W",      "note": "standard movement"},
  {"action": "move_back",     "input": "S",      "note": "standard movement"},
  {"action": "move_left",     "input": "A",      "note": "standard movement"},
  {"action": "move_right",    "input": "D",      "note": "standard movement"},
  {"action": "run",           "input": "Shift",  "note": "run while held"},
  {"action": "jump",          "input": "Space",  "note": "standard"},
  {"action": "attack_basic",  "input": "Mouse1", "note": "left click, primary action"},
  {"action": "skill_1",       "input": "1",      "note": ""},
  {"action": "skill_2",       "input": "2",      "note": ""},
  {"action": "skill_3",       "input": "3",      "note": ""},
  {"action": "skill_4",       "input": "4",      "note": ""},
  {"action": "skill_5",       "input": "5",      "note": ""},
  {"action": "skill_6",       "input": "6",      "note": ""},
  {"action": "skill_7",       "input": "7",      "note": ""},
  {"action": "skill_8",       "input": "8",      "note": ""},
  {"action": "inventory",     "input": "I",      "note": "Tab is ceded to target switching"},
  {"action": "map",           "input": "M",      "note": ""},
  {"action": "character",     "input": "C",      "note": ""},
  {"action": "menu",          "input": "Escape", "note": "standard"},
  {"action": "cam_lock",      "input": "Mouse3", "note": "wheel click toggles camera lock"},
  {"action": "zoom_in",       "input": "WheelUp",   "note": ""},
  {"action": "zoom_out",      "input": "WheelDown", "note": ""},
  {"action": "context_menu",  "input": "Mouse2", "note": "right click, secondary action"},
  {"action": "auto_move",     "input": "Mouse2", "note": "double right-click for auto-move"}
]

The output looks clean. That is what makes it dangerous. Look at the bottom two lines. context_menu and auto_move are both assigned to Mouse2 (right click). Even though the constraints clearly said "do not assign two actions to a single key," the model overlapped them anyway, attaching the excuse "distinguished by double-click" in the note. This is why AI output must never go straight into the build. A person scanning the table easily misses the conflict between lines 22 and 23 out of 23, and the model rationalizes its own conflict.

So I hand verification to code, not eyes. I run a small lint that checks conflicts (duplicate inputs) and reachability (missing required actions, placement outside the two-thumb corners).

# input_lint.py — input mapping conflict and reachability check
import json, sys
from collections import defaultdict

REQUIRED = {"move_forward","move_back","move_left","move_right",
            "attack_basic","menu","inventory","map"}

def lint(mapping):
    errors, warns = [], []
    seen = defaultdict(list)
    for m in mapping:
        seen[m["input"]].append(m["action"])
    # 1) Conflict: two or more actions on the same input
    for inp, acts in seen.items():
        if len(acts) > 1:
            errors.append(f"CONFLICT  {inp} <- {', '.join(acts)}")
    # 2) Reachability: missing required actions
    actions = {m["action"] for m in mapping}
    for r in sorted(REQUIRED - actions):
        errors.append(f"MISSING   required action '{r}'")
    # 3) Warn on empty note (design intent not recorded)
    for m in mapping:
        if not m["note"].strip():
            warns.append(f"NO_NOTE   {m['action']} ({m['input']})")
    return errors, warns

data = json.load(open(sys.argv[1], encoding="utf-8"))
errs, warns = lint(data)
for e in errs:  print("[ERROR]", e)
for w in warns: print("[WARN] ", w)
print(f"\n=> {len(errs)} error(s), {len(warns)} warning(s)")
sys.exit(1 if errs else 0)

Saving the JSON above as claude_map.json and running the lint, the actual output is this.

[ERROR] CONFLICT  Mouse2 <- context_menu, auto_move
[WARN]  NO_NOTE   skill_1 (1)
[WARN]  NO_NOTE   skill_2 (2)
[WARN]  NO_NOTE   skill_3 (3)
... (same for skill_4~8)

=> 1 error(s), 8 warning(s)

The lint pinpointed the single conflict the human eye had missed. The reachability check passed (all 8 required actions present). The 8 empty note entries are warnings only and do not block the build, but they expose a debt: design intent left unrecorded. Now I send it back to the model with the grounds for rejection.

Human Rejection + Re-request

Rejected: the lint shows context_menu and auto_move overlapping on Mouse2. Double-click discrimination adds right-click latency and misfires in combat. Split auto_move onto a separate input. Also, the notes for skill_1\~8 are empty — fill in one line each on what category of skill the slot holds.

Claude's Revised Output (Excerpt: Conflict Resolution Only)

  {"action": "context_menu", "input": "Mouse2",      "note": "right click = secondary/context action only"},
  {"action": "auto_move",    "input": "Numpad0",     "note": "auto-move toggle, physically separated from combat keys"},
  ...
  {"action": "skill_1", "input": "1", "note": "primary melee skill"},
  {"action": "skill_8", "input": "8", "note": "emergency dodge/survival skill — at the pinky's reach limit, consider remapping to Q"}

The last line of the revised output is interesting. The model flagged a reachability issue on its own, self-reporting that "the 8 key is at the pinky's reach limit." That is exactly the same topic as the reachability verification we cover in the next section. Run the lint again and it passes with 0 error(s). The core is this: AI produces a 23-line draft fast, but the draft's legality is guaranteed by human-defined rules (the REQUIRED set, the definition of a conflict) and by code. The model proposes, the tool judges, the human decides.

14.3.4 Input Flow — From Mapping to Screen

Draw the path by which one physical input becomes a game action, and you can see where the lint above slots in.

flowchart TD
    A[Physical input
touch coordinates / key·mouse] --> B{Input classification} B -->|contact 200ms↓ & 5px↓| C[Tap / click] B -->|contact 200ms↑ or 5px↑| D[Drag] C --> E[Mapping table lookup] D --> E E --> F{Lint-passed
mapping?} F -->|conflict·omission| G[Build blocked
input_lint.py] F -->|valid| H[Game action dispatch] H --> I[Action execution] I --> J[Feedback output
visual+haptic/sound] G -.fix and resubmit.-> E classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef fail fill:#fee2e2,stroke:#dc2626,color:#7f1d1d; class B,E,F code; class G fail;

A physical input entering at the top left is first classified as a tap or a drag (the 200ms/5px criteria in the next section). The classified input then looks up the mapping table — and the diagram's key point is that the table must pass input_lint.py before it enters the build. With a conflict or an omission, it never reaches the dispatch stage; it gets blocked. Mapping verification should be finished before runtime, at the build gate.

14.3.5 Five Principles of Touch Design

Now suppose the mapping has passed; we design the surface where that mapping meets the finger.

Principle 1 — Minimum touch area. Apple HIG 44pt and Material 48dp are the floor. On an HD screen, sizing buttons around 100px (200px in 2x Retina environments) satisfies both standards at once. Drop below this and the intro's "three taps, two fires" shows up as a statistic.

Principle 2 — Thumb reach zones. For mobile MMORPGs, the landscape two-handed grip is the standard: pressable elements go in the two bottom corners, and consumables/slots in the bottom center (for the basis of the three-zone model, see §9.1). P0 actions (left = movement, right = attack and skills) stay inside the two bottom corners; information that is rarely watched goes to the top, beyond the reach limit. The key fact for input design: even combined, the two corners cover less than half the screen. The SVG below shows the two-thumb reach zones and the bottom-center slot bar in landscape mode.

Top = out of reach (info) Bottom center = consumables·quick slots Left thumb Right thumb Easy Easy Dark = P0 button placement / light = reach limit

The dark fan is where the thumb lands without strain; the light fan is the limit reached only by stretching the hand. If the "8 key at the pinky's limit" the model reported in 14.3.3 is the PC edition of the problem, the mobile edition is the mistake of placing a P0 button in this light zone.

Principle 3 — Avoiding occlusion. A finger does not cover only the contact point; the whole hand above it covers the screen. Tap a bottom-right skill and roughly a quarter of the bottom right goes invisible. So display the results of an action (damage numbers, state changes) in areas fingers do not touch. The hand gripping the left joystick intrudes on the character and minimap positions, so the minimap moves to the top right.

Principle 4 — Drag/tap distinction. Unlike a mouse, touch must infer the user's intent from time and distance. Unify on one criterion across the entire game — for example, contact within 200ms with movement within 5px is a tap; anything beyond is a drag. When these two numbers wobble, intent failures pile up: "I tried to tap and my character rolled." The branch point in the mermaid diagram above is exactly this judgment.

Principle 5 — Haptics. Vibration is the only channel that communicates without the screen being watched. But give every input a vibration and it becomes noise. Plain taps: none; skill use: short; risky actions like purchase confirmation: strong; enemy kills: subtle — keep it within 4–5 variants.

14.3.6 Five Principles of Mouse Design

The mouse enjoys three luxuries touch does not have: hover, multiple buttons, and cursor precision.

Principle 1 — Hover. A mouse can point without pressing. Rest it on a skill slot and a tooltip with name, cooldown, and description appears; click and the skill is used. Touch has no such in-between state, so hover is a channel through which PC can layer on more information. The caveat: information that depends on hover alone has nowhere to go in the mobile edition — something to be conscious of back at the 14.3.3 mapping stage.

Principle 2 — Multiple buttons. Left click is the primary action, right click is secondary/context, wheel click resets the camera, the wheel zooms. The conflict the lint caught above was precisely a case of stacking two actions on this right click. Trying to fill every button just because they exist is how conflicts get made.

Principle 3 — Keyboard conventions. ESC = menu, M = map, 1–8 = skills, WASD = movement, Shift = run, Space = jump. Users should be able to guess without learning. Keys that deviate from convention get their justification written in note — the decision in 14.3.3 to cede Tab to target switching instead of inventory is one example.

Principle 4 — Camera control. Rotate the camera with mouse drag, but toggle clearly between a game mode that locks the cursor to the screen and a UI mode that releases it. When this toggle is ambiguous, you get the confusion of closing a menu only to have the cursor vanish.

Principle 5 — Macro and automation allowances. How far to allow auto-attack and auto-move is a question of game identity. Loosen it too much and PC becomes a screen full of macros; ban it outright and the entry barrier rises for users coming over from mobile. The answer is picking the point on the spectrum that fits the game's color, not either extreme.

14.3.7 Common to Both — Unified Input Feedback

The principles differ per platform, but the "feel" a user gets from the same action must stay the same when the platform changes. A user moving from mobile to PC should not have to relearn what a glowing button means.

Situation Touch Mouse
Input recognized Button glow + short haptic Button glow + click sound
Input failed Button shake + haptic Button shake + warning sound
Cooldown in progress Radial gauge Radial gauge
Ready again Glow + haptic Glow + sound

The visual channel (glow, shake, gauge) is identical on both sides; only the secondary channel splits into haptic vs. sound per platform. This consistency cuts the learning cost for multi-platform users in half.

14.3.8 Common Failures and Fixes

Pattern Fix
Buttons below the standard floor (44pt/48dp) Enforce around 100px
Results displayed in the finger-occluded zone Move to a non-occluded zone
Haptic overuse Keep within 4–5 variants
Two actions stacked on right click Block the conflict with lint, then split
Hover-only information carried to mobile as is On mobile, substitute tap/long-press channels
Hard-locked key mapping Allow user customization
Forcing identical mapping on both platforms Natural mapping per platform

The fourth row of this table is the conclusion of the 14.3.3 worked transcript. Scanning the table by eye, the right-click conflict is missed almost every time; hang the lint on the build gate and it is caught almost every time.


Key Takeaways

Next Chapter Preview


Try It Yourself (setup → prompt → verify)

  1. setup — Collect your action list and constraints in one file (movement, attack, skills, UI, camera). Put the input_lint.py above in your project. Replace the REQUIRED set with your own game's required actions.
  2. prompt — Use the full prompt from 14.3.3 as is, swapping in only your action list. Pin the output format to a JSON array only.
  3. verify — Run the returned JSON through python input_lint.py claude_map.json. Until ERROR reaches 0, re-request with explicit grounds for rejection (conflicting inputs, missing actions). Log WARN (empty notes) separately as unrecorded-design-intent debt.

Solo Scale-Down

If you are building a small game alone, shrink the tooling. With fewer than 10 actions, a 20-line lint that keeps only two checks — the REQUIRED set and "duplicate input" detection — is enough. Get the mapping from AI, filter only the conflicts through this mini lint, then press the buttons on a real device to see whether your thumb (or pinky) reaches. The model proposes, code judges conflicts, your own hand judges reach — hold to these three and it works at any scale.

Part 15 · Live Ops

15.1 Live Ops Overview — AI Combines Event Candidates, the Rulebook Filters Them, and a Human Chooses

Primary audience: a game designer taking charge of live operations (live ops) for the first time after launch, on a mid-size team (10–50 people) Scaled-down version for solo/hobbyist readers: §15.1.7, "If you're solo, just this much"

Premise: I ran live ops on a globally launched mobile MMORPG, including its P2E (play-to-earn) economy, and I write this chapter by combining that experience with the pre-launch AI workflow on my current project. The worked transcript is the result of actually running the "input → AI combination → rulebook verification → human selection" pattern once on live ops formats. Estimates and observations are labeled as estimates and observations, and no made-up KPI table is included.

The office on the morning after launch is not the office from before launch. The milestone is over, yet the work does not shrink — only its unit gets smaller. Schedules that ran in quarters split into weeks, days, and hours. And every week, the same question returns to the meeting room: "What event are we running this weekend?"

If that question restarts from a blank page every week, the live ops team burns out fast. This chapter is about taking that question off the blank page. The core is twofold. First, instead of squeezing out events and seasons from scratch every time, accumulate them as a library of proven formats. Second, hand the tedious drafting — "combine those formats into 5 candidates for next week" — to AI, and let the human decide only which of the candidates that passed rulebook verification to adopt. Building from zero and picking one of five are very different workloads.


15.1.1 Live Ops Is a Loop, Not a Feeling

Plenty of books teach the standard live ops cycle as a table to memorize: report on Monday, prepare on Tuesday and Wednesday, deploy on Friday. All true — but memorizing the table never shows you how the decision that comes back every week, "this week's event," actually gets made. The essence of live ops is not a schedule but a closed loop — candidates appear, pass verification, a human picks, the build ships, and player data comes back as the input for the next candidates. One full turn.

On top of this loop, the four tracks of live ops — content, events, balance, and customer support (CS) — each run at their own speed. Content runs monthly to quarterly, events weekly to monthly, balance weekly to biweekly, and CS daily and hourly. When the four tracks run separately, the same player data produces a different decision every week. So the goal of this chapter is to tie the four tracks into a single loop, and to shape one cell of that loop — event candidate generation — into a form AI can run.

flowchart TD
    A["Input
KPI trends · player segments
· past event results"] --> B["AI combination
season rules × event template
library → 5 candidates"] B --> C{"Rulebook verification
inflation cap · purpose conflicts
· reward range · duration"} C -->|violation alert| D["Re-request
(replace violating candidates)"] D --> B C -->|pass| E["Human selection
director adopts/rejects"] E --> F["Build & deploy
(irreversible: season start · announcement)"] F --> G["Player data & feedback
(input for the next loop)"] G --> A classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; class A,G data; class B ai; class C,F code; class E human;

Human hands touch this loop in exactly two places: at the top, where the input (KPIs, segments, past results) is fed in clean, and at the spot where someone decides which of the verified candidates goes live. The tedious work in between — squeezing out 5 combinations and filtering rule violations — is run by AI and the rulebook. And the single line at the bottom — the player data created by the shipped event flows back in as input — is what makes this loop live ops. Pre-launch design ships once and is done; in live ops, the result becomes the next input.

The two libraries that feed this loop (season rules and event templates) are detailed in §15.2, and the last cell (automatic classification of player feedback) in §15.3. This chapter focuses on going around the loop once, all the way.


15.1.2 [Worked Transcript] Combining 5 Event Candidates → Rulebook Verification → Human Selection

Here is one full cycle, end to end, of how this actually runs. Below is a reproduction of a session in which I took the "library combination → rulebook verification → human selection" pattern — verified in my pre-launch content tooling — and ran it once on live ops formats (season rules + event templates). The input prompts can be copied and used as they are; the outputs are a reconstruction of that session.

Step 1 — Input: Hand Over the Library and the Current State as They Are

First, put the two ingredients for combination in a form a machine can read: the event template library (proven formats) and the season rule library — plus this week's current state (KPIs and segments). The libraries are built once and reused every week.

# event_templates.yaml — library of proven event templates (excerpt, 4 of 9)
- id: tpl_attendance      # daily login rewards
  purpose: [acquisition, reactivation]
  recommended_duration: 7–14 days
  reward_tier: low–mid
- id: tpl_coop_raid       # co-op raid
  purpose: [engagement, community]
  recommended_duration: 3–7 days
  reward_tier: mid–high
- id: tpl_pvp_season      # competitive season
  purpose: [community, engagement]
  recommended_duration: 14–28 days
  reward_tier: high
- id: tpl_limited_package # limited-time package
  purpose: [revenue]
  recommended_duration: 3–7 days
  reward_tier: high (purchase-linked)

# season_rules.yaml — season rule fragments (excerpt)
season_inflation_cap: high-tier reward events ≤ 3 per quarter
purpose_conflict_rule: no two [revenue]-purpose events in the same week
overlap_rule: no two high-reward events at the same time (fatigue, inflation)

# current_state.yaml — this week's state
week: 2026-W23
revenue_events_last_2_weeks: 1 (quarterly high-tier total: 2)
DAU_trend: gradual decline (-6% over the last 4 weeks; 'caution' zone by industry observation)
key_segment: rising share of lapsed players likely to return
upcoming_external_events: none

Step 2 — Prompt: Ask for Combinations, but Enforce Format and Rules

Using the attached template, season-rule, and current-state YAML files, combine exactly 5 event candidates for next week.
Do not invent new mechanics — use combinations of the attached templates only. Flag any season-rule violations yourself,
and add one line per candidate explaining why it fits the current signals (DAU decline, lapsed-player return).
Spread the purposes so they don't all skew toward revenue, and flag anything ambiguous and hand it to me.
Table only; columns: candidate, template combo, primary purpose, duration, rule violation, rationale, ambiguous.

Notice that not one of the five lines of rules says what to make. Every one of them is a constraint — "answer within this frame, based on this evidence." Give AI free rein and it spits out the average event of a generic RPG. Restrict the materials to the library (rule 1), require a self-check against the season rules (rule 2), require a rationale tied to the current signals (rule 3), spread the purposes (rule 4), and leave an escape hatch — "if it's ambiguous, hand it to the human" (rule 5). Together, the five lines define not the shape of the candidates but their qualifications.

Step 3 — Claude's Output (Reconstruction, Excerpt)

Below is the raw output, unsummarized. Rule violations and ambiguity flags are mixed in.

| Cand. | Template combo               | Primary purpose         | Duration | Rule violation | Rationale                                              | Ambig. |
|-------|------------------------------|-------------------------|----------|----------------|--------------------------------------------------------|--------|
| C1    | tpl_attendance               | reactivation            | 10 days  | none           | directly targets rising lapsed-return share; reward 'mid' | -      |
| C2    | tpl_coop_raid                | engagement              | 5 days   | none           | login incentive during DAU decline; stirs community      | -      |
| C3    | tpl_attendance+tpl_coop_raid | reactivation+engagement | 7 days   | none           | logins pull returners + raid retains them; combo synergy | -      |
| C4    | tpl_pvp_season               | community               | 21 days  | none           | long season lifts the retention curve itself             | ★      |
| C5    | tpl_limited_package          | revenue                 | 5 days   | ★violation     | quarter already at 2 high-tier → adding nears/exceeds cap | -      |

Rule check results:
- C5: risk of violating season_inflation_cap. This quarter already has 2
  high-tier reward events. Adding C5 makes 3, right at the cap. A revenue
  push during a DAU decline can add to player fatigue — it also runs
  against the current signals.
- C4: no rule violation, but a 21-day PvP season occupies 3 weeks, not
  just this one. Whether it fits as a "this week" candidate is an ops
  policy call → flagged as ambiguous.
- C1, C2, and C3 pass the rules and align with the current signals
  (lapsed players, DAU).

The most valuable part of this output is not the table but the three rule-check lines at the bottom. The AI caught and flagged C5's inflation violation on its own, and handed C4 to the human with "no violation, but judgment may differ." A good prompt lets AI report its own limits and pass the ball to the human.

Step 4 — Verification and Selection (the Human's Part)

You do not accept this output as is. Run it through the rulebook once more, and then a human picks. In this session, two calls actually split.

First, C5 is rejected. The AI had already flagged the inflation violation, and the rulebook code (§15.1.3) returned the same verdict. It hits the quarterly high-tier cap, and a revenue push during a decline in daily active users (DAU) runs against the current signals. There is nothing to debate. Out it goes.

Next is C4 (the 21-day PvP season) — the one the AI flagged as ambiguous. No rule is violated, but this is not a "this week's event"; it is a "this season" decision. It is not something to settle on the spot in a one-week loop — it belongs in the integrated season meeting. So it is held out of this week's candidates and pulled aside as a season calendar agenda item.

From the remaining C1, C2, and C3, the director picks. The best fit for the current signals (a rising share of returning lapsed players + a gradual DAU decline) was C3 (login rewards + co-op raid combined). Login rewards pull lapsed players back in, and the raid holds on to the players it pulled in — that combined synergy matched this week's signals. C1 and C2 stay in the candidate pool for next week.

One candidate was not finished here. After deciding to adopt C3, it turned out the last day of the 7-day run overlapped with an upcoming scheduled maintenance day. So one re-request goes around.

We adopt C3. However, the last day of the 7-day run overlaps with a scheduled maintenance day.
Adjust the duration and propose again, so that maintenance does not cut off participation at the end of the event.
Keep the total reward amount the same and only move the schedule earlier.

The AI answered again with the start date moved up a day so the event ends before maintenance, and that adjustment passed the rulebook. Input → AI combination → rulebook verification → human selection → schedule re-adjustment: the cycle closes here.

This one full turn is the standard of Show this entire book holds itself to. Unless you watch, at least once and all the way through, what the AI combines, what the rulebook filters, and what the human picks and rejects, the sentence "we generate event candidates with AI" is hollow.


15.1.3 The Rulebook as Code — Automatic Candidate Verification

Check by eye every week whether the candidates respect the season rules, and you will miss things — again. Of the three rules in §15.1.2, whatever can be judged numerically should be reviewed by code. Humans spend their time only on the "ambiguous" and the "selection" that code cannot catch.

# event_lint.py — verifies next week's event candidates (skeleton)
# Input: candidate list combined by AI + season rules + quarterly running state
# Output: list of rule violations (alerts, not auto-rejection)

def lint(candidates, season, quarter_state):
    issues = []
    high_used = quarter_state["high_reward_count"]  # high-tier count so far this quarter
    for c in candidates:
        # Rule A: inflation cap (high tier ≤ 3 per quarter)
        if c["reward_tier"] == "high" and high_used + 1 > season["inflation_cap"]:
            issues.append(f"[A] {c['id']}: adding a high-tier event exceeds the quarterly cap "
                          f"of {season['inflation_cap']} (currently {high_used})")
        # Rule B: no two [revenue]-purpose events in the same week
    sales = [c for c in candidates if "revenue" in c["purpose"]]
    if len(sales) > 1:
        issues.append(f"[B] {len(sales)} [revenue]-purpose candidates at once → limit to 1")
        # Rule C: purpose skew (if one purpose is a majority of the 5, spread is insufficient)
    from collections import Counter
    top = Counter(c["primary_purpose"] for c in candidates).most_common(1)[0]
    if top[1] > len(candidates) // 2:
        issues.append(f"[C] purpose '{top[0]}' appears {top[1]} times — skewed (insufficient spread)")
    return issues

This code settles the meeting-room back-and-forth — "isn't this reward too generous?" — with a single line of numbers. When the code prints [A] tpl_limited_package: adding a high-tier event exceeds the quarterly cap of 3 (currently 2), there is nothing to debate. You drop it. This is the lint gate from §14.1 (mobile HUD) carried over to the live ops level — the division of labor where code catches what determinism can catch and humans take what requires judgment holds in operations just the same.

One thing is different, though. Even when this lint finds a violation, it does not automatically discard the candidate. It only raises an alert. It is the same design we saw in the city generator in §6.2. Attach auto-rejecting verification, and the machine also kills intended variations — say, a campaign decision to run a revenue event knowing full well where the quarterly cap stands. The machine surfaces suspect candidates; whether they live or die is the director's call. Rejecting C5 in §15.1.2 was not the lint killing it, either — it was a human decision made after seeing the lint's alert.


15.1.4 Before and After Launch — What Changes

The loop above differs decisively from the pre-launch design loop on two points. Rather than listing a table, I will pin down exactly these two.

First, the result becomes the next input. Before launch, a game design document flows one way: you write it, and it flows through to the build. In live ops, the player data this week's event creates (participation, churn, revenue, feedback) comes back as the input (current_state.yaml) for next week's candidate combination. The bottom arrow in the §15.1.1 loop is that return. So the KPI of live ops is not "get it right once" but "adjust to the signals every week."

Second, experiments get cheaper, but the irreversible points get sharper. Before launch, a single decision shaped a quarter; live, you run a one-week event, and if it misses, you change it next week. Rollback-friendly experiments multiply. But a season start and an event announcement are irreversible. The principle from §5.4.5 — "recording and casting = irreversible steps" — applies as is. Season rules and rewards that players have already seen leave a mark on community perception even if you "cancel" them. So every verification in the §15.1.1 loop — AI combination, rulebook, human selection — must finish in the reversible stage, before entering the irreversible cell of build and announcement. Holding C4 (the 21-day season) out of this week's snap decision and sending it up to the season meeting follows the same principle — the bigger the irreversible point of a decision, the longer the reversible review it gets.

These two points make live ops a different job from pre-launch design. The rest — time units shifting from quarters to weeks, feedback shifting from beta tests to real time — derives from these two axes.


15.1.5 From Conservative to Progressive Application

The worked transcript in §15.1.2 is a scene from the progressive application: AI combined the candidates, and the human decided the adoption. But not every team gets there from day one. There are stages.

In the conservative application, humans originate the candidates. The ops team designs the events directly in the Monday meeting, writes the season rules by hand, and classifies player feedback manually. Automation covers only measurement (the KPI dashboard) and regression checks (build verification). By industry observation, most live MMORPG operations today are close to this stage.

In the progressive application, AI drafts even the "event candidate proposals" and the "feedback classification." §15.1.2 is a scene of the former; the latter (automatic feedback clustering) appears in §15.3. Human decisions narrow to meta-decisions: which candidate to adopt, and how to take the feedback AI has classified.

For the progressive application to take root, three things must be in place: a library in which event templates and season rules are separated and accumulated as recombinable units (the event_templates.yaml in §15.1.2 is its seed), a candidate generator that takes the current signals as input and produces draft candidates (the prompt in §15.1.2), and clustering that automatically classifies incoming feedback (§15.3). That these three share the same skeleton as §5.3.12 (the world Behavior Tree (BT) and the quest cloud) and §8.1.8 (progressive balancing) is this book's consistent message — different domains, same structure: accumulate verified pieces as a library, let AI propose combinations, and let a human adopt.

Let me make one thing clear here. Ideas like the library, the candidate generator, and clustering were theoretically possible even in the 2010s. What blocked them was that AI could not write the natural language players actually read — event announcements, rule explanations — and could not summarize and classify hundreds to thousands of feedback items a day in natural language. After the advances in large language models (LLMs) from 2023 onward, those two walls got lower, and a large part of progressive operations that had existed only on paper moved into the realm of the feasible.


15.1.6 Common Failures

Pattern Why It Fails Remedy
Designing events from a blank page every week The ops team burns out fast; candidate quality swings with the team's condition Accumulate an event template library (§15.1.2)
Wholesale delegation — "AI, make me an event" Without a library and rules, you get the average of generic RPGs Restrict the materials + force a season-rule self-check (§15.1.2)
Reviewing candidates by eye only Inflation and purpose skew slip through every week Verify automatically with event_lint.py (§15.1.3)
Making the lint auto-reject The machine kills intended campaign decisions too Alerts only; adoption is the director's call (§15.1.3)
Settling irreversible decisions on the spot in the weekly loop Rolling back after a season announcement leaves a mark on the community Split big decisions out to the season meeting (§15.1.4)
Chasing a single KPI (DAU or revenue) Player fatigue accumulates; candidates get adopted against the signals Feed multi-axis signals into current_state (§15.1.2)

15.1.7 Try It Yourself — One Step You Can Take Today

Try just one step, in setup → prompt → verify order.

If you're solo, just this much: You don't need the library YAML or the lint code. Recall just 5–6 events from the last quarter of a game you love and write them down in three columns — purpose, duration, reward. That alone shows you the game wasn't squeezing events out of a blank page every week; it was rotating formats. That table is your first template library.

If you're on a team, start with this one step: gather the events from the last one or two quarters, normalize them into event_templates.yaml (proven formats only), and put the three season rules into code first with event_lint.py. With a library and rules in place, you can measure AI-combined candidates and human drafts with the same ruler.


Key Takeaways

Next Chapter Preview

15.2 Event and Season Ops — From One Template to Ten Variation Candidates, Only the Review Is Human

Primary audience: MMORPG designers responsible for live ops (mid-sized teams of 10–50) Scaled-down version for solo/hobbyist readers: §15.2.9, "If You Work Solo, Do Just This Much"

I think back to the Monday meetings on a live game in its fourth year of service. Every week, the question of what event to run next week started from a blank page. Someone would say, "How about the attendance event from last time, with the rewards bumped up a bit?" Someone else would say, "We did that two months ago." And how much to bump the rewards was, once again, decided by gut. After the meeting, one live-ops designer spent half a day filling out the event form from scratch. Every week, from a blank page, half a day.

The problem was not a shortage of ideas. The ops team already carried a few proven event skeletons in their heads: attendance, co-op, competition, comeback. Swap a new theme and new rewards into one of those skeletons and you have a one-week event. But because that swapping was done by hand and by gut every time, it was slow and the results wobbled.

This chapter covers how to hand that swapping over to AI. Two things matter. First, encode the proven event skeletons as variation-ready template YAML. Second, give AI the tedious job of pulling multiple candidates for next week out of the template, while humans enforce reward ranges and overlap in code, then review only the tone. The general theory of event design (attendance events help new-user inflow, co-op events help engagement, and so on) is already well covered in other books, so this chapter focuses only on running that knowledge as an AI workflow.

A frank note on my ops experience Direct responsibility for post-launch live ops over one-to-two-year stretches covers only part of my career. The workflow in this chapter is my production-and-review tooling (content and HUD) carried over into events, and the effect figures are flagged in the text, case by case, as industry observation plus author's estimate. The tool structure (template YAML, lint, review gate) shares its skeleton with the content production tools I actually operate.


15.2.1 Humans Write the Template and Do the Final Review

The full flow of event production has four stages. The key point: stage 1 (template) and stage 3 (lint) are deterministic, and only stage 2 is AI. It is the same division of labor we saw in content production (§6.2) and HUD compression (§14.1). With the rulebook holding down both the input end and the verification end, the AI sandwiched in the middle can produce a slightly different variation every time without the reward balance or the schedule wobbling.

flowchart TB
    A["Input: event template YAML
(proven skeletons — attendance, co-op, competition, comeback)
+ this quarter's theme, banned rewards, calendar slots"] A --> B["Stage 2 AI: generate variation candidates
same skeleton × different theme, rewards, duration
→ 5~10 candidates (draft forms)"] B --> C{"Stage 3 deterministic: event_lint.py
reward range, inflation cap, schedule overlap,
same skeleton repeated within the last N weeks"} C -->|WARN on violation| D["Ops review gate
(adopt, reject, fine-tune)"] C -->|pass| D D -->|re-request| B D -->|adopt| E["Into the build → announcement
(irreversible gate)"] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; class A data; class B ai; class C,E code; class D human;

Human hands touch this diagram in exactly two places. At the top, where the template and this quarter's constraints get entered cleanly. At the bottom, where someone judges what lint cannot catch — "does this theme fit the mood of our game right now?" The tedious candidate generation and reward arithmetic in between is run by the template, the AI, and the lint.

The decisive design choice is that when lint (stage 3) finds a violation, it does not discard the candidate automatically — it only raises a WARN to the ops gate (stage 4). The reason comes in §15.2.5. And the final arrow (the announcement) being irreversible is what sets live ops apart from other production work. A city NPC you dislike can simply be scrapped before the build; an event announced to players costs community trust to roll back (§15.2.7).


15.2.2 Input — The Event Template YAML

Pin the ops team's proven skeletons down into a fixed form. Left as a free-form design doc, the AI does not know what it is supposed to vary. Only when the slots are separated does "swap only this slot" become a workable instruction.

# event_templates/coop_raid.yaml — co-op raid skeleton (proven, run 4 times)
template_id: coop_raid
purpose: [existing_player_activation, community]      # 1~2 only. Never chase all four at once
core_loop: the whole server accumulates contribution during the event → server-wide rewards unlock tier by tier
duration_range: [5, 10]              # days. Past 10 days, fatigue builds up
slots:                               # ← the fields AI varies. The skeleton is fixed
  theme: { type: free, 제약: quarter_theme_compliance }
  boss_or_target: { type: free, 제약: reuse_existing_boss_assets_first }
  reward_tiers: { type: reward_list, count: 3~5, 제약: see reward_policy }
reward_policy:                       # ← the fields lint reads. No variation allowed
  강화석_per_event_max: 30           # payout cap per single event
  골드_per_event_max: 50000
  한정코스튬: allowed (permanent ownership, zero economy impact)
  현금성재화_직접지급: prohibited
inflation_guard:
  강화석_분기_누적상한: 90           # summed across all events in the quarter
post_event_kpi:                      # ← auto-measured post-event slots
  - participation rate (joined at least once, out of players shown the event)
  - enhancement stone price movement (30 days post-event, target ±10%)
  - weekday playtime after the event (dependency signal)

The most important separation is between slots (which AI varies) and reward_policy (which lint reads and AI may not touch). Theme and boss can differ every time, but the enhancement stone payout cap is a line drawn by the game economy. If the AI pulls a different number for that line on every call, inflation starts right there. So the AI proposes the reward items, while lint enforces that the reward amounts move only within the policy range.

The same folder holds daily_attendance.yaml (attendance), pvp_ladder.yaml (competition), and comeback.yaml (comeback) in the same format. These four files are the entire input pool for this quarter's candidate generation.


15.2.3 [Worked Transcript] One Template → Variation Candidates

Here is one full cycle, end to end, exactly as it runs. The input prompt can be copied as is; the output is a reconstruction of an actual production session.

Step 1 — The Prompt: Ask for Variation, Enforce the Skeleton and the Policy

Using the attached coop_raid.yaml (a co-op raid skeleton proven over 4 runs), give me exactly 5 variation candidates for next week (W2).
This quarter's theme is "summer — water, festival, heat".
Do not touch core_loop at all; change only the slots (theme, boss, rewards).
Keep rewards within the reward_policy caps, and attach one line per candidate explaining why this theme and this reward.
The last 3 weeks were attendance, PvP ladder, and co-op raid, so tag any candidate that runs a co-op raid again with [repeat-risk],
and tag anything you are not sure about with [unsure] and hand it to me. Table only; columns: candidate name, theme, reward tier summary, rationale, flags.

Of these six lines of rules, not a single line specifies a theme or a reward. Every line is a fence: change only this much, do not cross that line, report when confused. Hand it a loose "make me a summer event" and you get the average of every mobile RPG out there. Variation scope (slots only), reward caps (reward_policy), repetition alert (the last 3 weeks' history), and "when in doubt, hand it to a human" — these fences keep the skeleton intact while setting the variation free.

Step 2 — Claude's Output (Reconstructed from an Actual Session, Excerpt)

| Candidate            | Theme                      | Reward tier summary                            | Rationale                                                    | Flags         |
|----------------------|----------------------------|------------------------------------------------|--------------------------------------------------------------|---------------|
| Flood Levee          | hold back the river flood  | 25 enhancement stones / 40,000 gold / costume  | co-op = levee-building metaphor, maps to cumulative contribution |           |
| Floating Lanterns    | summer water festival      | 20 enhancement stones / 30,000 gold / title    | festival tone, mood refresh rather than competition          |               |
| Heatwave Boss Battle | subdue the avatar of heat  | 30 enhancement stones / 50,000 gold / costume  | reuses existing boss assets, rewards at the cap              | [repeat-risk] |
| Ocean Current Race   | ride the currents, compete | 28 enhancement stones / 45,000 gold / mount    | co-op + competition mix                                      | [unsure]      |
| Koi Collection       | collection-based co-op     | 22 enhancement stones / 35,000 gold / title    | low entry barrier for casual players                         |               |

[repeat-risk] Heatwave Boss Battle: core_loop is effectively identical to the W-3 co-op raid.
  Same skeleton repeated at a 3-week interval → player fatigue risk. A different skeleton (attendance/collection) is recommended.
[unsure] Ocean Current Race: mixes cumulative co-op and individual competition in one event.
  The purpose slot grows to [activation + competition]. Possible conflict with the 1~2 purpose rule.

The most valuable part of this output is not the table but the two notes at the bottom — the place where the AI reports its own limits ("Heatwave Boss Battle shares its skeleton with the event three weeks ago," "Ocean Current Race has grown a second purpose") and hands the call to a human. A good prompt is one that lets the AI say, "I am not confident about this one."

Now lint runs over this batch of candidates.


15.2.4 Stage 3 Lint — Reward Ranges and Overlap, Enforced in Code

Checking by eye, every time, whether a candidate respects the reward policy and avoids schedule overlap is exactly how things get missed — again. Anything that can be judged from reward_policy, inflation_guard, and the calendar should be reviewed by code. People spend their time only on the tone and fun judgments code cannot make.

# event_lint.py — validates event variation candidates (skeleton)
# Input: candidate list proposed by AI + template policy + quarter calendar
# Output: list of WARNs (not auto-discarded — escalated to the ops gate)

def lint(candidates, policy, quarter_ledger, recent_weeks):
    warns = []
    stone_used = sum(quarter_ledger.강화석)   # cumulative payout already granted this quarter
    for c in candidates:
        # A: per-event reward cap (policy)
        if c.강화석 > policy["강화석_per_event_max"]:
            warns.append(f"[A] {c.name}: enhancement stones {c.강화석} > cap "
                         f"{policy['강화석_per_event_max']} (per-event excess)")
        # B: quarterly inflation cumulative cap
        if stone_used + c.강화석 > policy["강화석_분기_누적상한"]:
            warns.append(f"[B] {c.name}: quarter total {stone_used + c.강화석} > "
                         f"{policy['강화석_분기_누적상한']} (inflation cap)")
        # C: same skeleton repeated within the last N weeks
        if c.template_id in recent_weeks[-2:]:
            warns.append(f"[C] {c.name}: {c.template_id} skeleton ran within the last 2 weeks (repeat)")
        # D: calendar slot conflict (another major event in the same week)
        if quarter_ledger.slot_taken(c.week):
            warns.append(f"[D] {c.name}: W{c.week} slot already holds a major event")
    return warns

Feed the five candidates from the worked transcript above into this code, and this comes out.

[PASS] Flood Levee: enhancement stones 25 ≤ 30, quarter total 65+25=90 ≤ 90 (boundary reached)
[WARN] [C] Heatwave Boss Battle: coop_raid skeleton ran within the last 2 weeks (W-3) (repeat)
[WARN] [B] Ocean Current Race: quarter total 65+28=93 > 90 (inflation cap exceeded)
[PASS] Floating Lanterns: enhancement stones 20 ≤ 30, quarter total 65+20=85 ≤ 90
[PASS] Koi Collection: enhancement stones 22 ≤ 30, quarter total 65+22=87 ≤ 90

The interesting one here is Ocean Current Race. The AI tagged it [unsure] over the purpose conflict, but lint caught it for an entirely different reason — exceeding the quarterly cumulative inflation cap. Add its 28 enhancement stones and the quarter total reaches 93, past the policy's 90. The code caught arithmetic the AI missed. Conversely, on Heatwave Boss Battle, the AI's [repeat-risk] and lint's [C] pointed at the same thing. Human, AI, and code each filter with a different net.

Thanks to these 30 lines, "isn't this reward a bit rich?" no longer ends as gut versus gut. When the code prints [B] quarter total 93 > 90, there is nothing to debate. Lower the reward or swap the candidate.


15.2.5 One Full Cycle — Review, Rejection, Re-Request

Writing "the ops team reviews it" in the abstract tells you nothing about what this gate actually filters. Let us follow one cycle to the end: after lint passes, what does a human kill, and what does a human keep?

[Stage 4 Ops Review — Verdicts]

The live-ops designer handled the five candidates like this.

The heart of this gate is that a human shook the lint-passing Flood Levee out of the top spot. The code passed 90 ≤ 90. No policy violation. But the live-ops designer was looking at the reward rhythm of the entire quarter. Lint sees the legality of one event; a human sees all the way to the season finale at quarter's end. So a re-request goes out.

Remake the Flood Levee variation with its reward lowered from 25 enhancement stones to 18.
Reason: we must keep 12 enhancement stones of headroom for the season-finale push in the last week of June.
To offset the weaker reward appeal, restructure reward_tiers to shore up
perceived value with limited costumes and titles instead of enhancement stones.

The AI came back with a candidate that lowered the enhancement stones to 18 and expanded the limited costumes to 2 (permanent-ownership rewards with zero economy impact). Re-running lint gave quarter total 65+18=83 ≤ 90, leaving headroom of 7 for the season finale. The cycle — input → candidate generation → lint → review → rejection → re-request — closes here.

This one loop is the Show standard of this entire book. Unless you watch, at least once and end to end, what the tool emits, what gets caught, and what a human kills, the sentence "we mass-produced events with AI" is hollow.

This cycle is also why the lint is not an auto-discarder. Had lint silently dropped the [B] violation, the ops team would have lost the chance to learn Ocean Current Race's real problem (the purpose conflict), and there would have been no place to shake a candidate like Flood Leveelegal, yet risky for the quarter's rhythm. The machine nominates the suspects; humans decide adoption and rejection.


15.2.6 Seasons — A Bigger Rhythm, the Same Separation

If events run on a weekly-to-monthly rhythm, seasons run on a quarterly one. The operating method is the same. Separate a season's proven elements into slots, and each quarter you swap only the theme.

Season slot Varies (AI and humans) Fixed (policy and lint)
Season theme summer, winter, new year (free)
Season pass reward track reward items per tier number of tiers, completion difficulty, reward caps
Season PvP ranking ranking reward items reward inflation cap
Meta shuffle new characters, balance change-magnitude guardrails (§8.1)

For the season pass, the key figure humans pin down as policy is the completion rate target. A commonly cited industry benchmark is to tune difficulty so that around 70% of active users reach the final tier (author's estimate — it varies by game, so read it as a direction, not an absolute: below 30% means frustration, above 90% means no sense of challenge). Once this target sits in a slot, you can force the AI to produce an "expected completion rate" alongside any season pass variation it proposes.

Events and seasons stop colliding only when the quarterly calendar is visible at a glance. It is closer to the ops team's shared desk calendar than to a chart. Conflicts shrink when everyone is looking at the same picture.

Q2 combined calendar — one season (quarter), events by week Season "Summer Festival" (50-tier season pass · PvP ranking) — April–June, all quarter April May June W1 Attendance W2 Co-op (Levee) W3 PvP Ladder W4 Collection Co-op W5 Comeback W6 Season Finale Quarterly enhancement stone inflation total (cap 90) current total 83 / cap 90 (headroom 7 = reserved for the season-finale push) Attendance/Collection Co-op/Comeback Competition (PvP) Season event

This one picture explains the §15.2.5 judgment visually. Color is skeleton type. If the same color appears twice within 2–3 weeks, the §15.2.4 lint [C] cries out. And the inflation gauge at the bottom sits just short of the red line (cap 90), with barely 7 of headroom left for the June season finale (W6) — the same 7 secured by lowering the Flood Levee reward to 18.


15.2.7 The Irreversible Gate — Finish Every Review Before the Announcement

One thing decisively separates live ops from city NPCs (§6.2) or the HUD (§14.1). An announcement cannot be undone. An NPC whose tone is off can simply be scrapped before the build, and players never know it existed. But once an event has been announced, its rewards, duration, and rules live on in the community. Saying "the event rewards were too generous, so we are clawing them back" after launch carries an irreversible cost.

flowchart LR
    A["Template variation candidates"] -->|reversible| B["Lint review"]
    B -->|reversible| C["Ops review, rejection, re-request"]
    C -->|reversible| D["Build review
(release can be held)"] D ==>|irreversible gate| E["Event announcement and start"] E -.->|recovery is costly| F["Only late-stage fine-tuning possible
(reward clawbacks and duration changes cost trust)"] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef fail fill:#fee2e2,stroke:#dc2626,color:#7f1d1d; class A data; class B code; class C,D human; class F fail;

The principle that runs through this whole book (§5.4.5 voice recording, §8.1 live builds, the same message as Part 12's final render) holds in live ops too. Every review — reward range, inflation cap, schedule conflicts, tone — must finish in the reversible stages before the announcement. That is why the entire production–lint–review–re-request cycle of §15.2.3\~5 turns to the left of this irreversible gate. Once past the gate, all that remains is the late-stage fine-tuning of §15.2.8, and even that nibbles away at player trust.


15.2.8 Signals and Prescriptions During Operation

KPIs are still watched after the announcement. But unlike pre-announcement review, the only lever left here is late-stage fine-tuning. The signals are measured automatically; the prescriptions are human decisions.

Signal (auto-measured) Prescription (human decision)
Participation below 50% Slightly strengthen late-stage rewards or extend the duration by +2 days (within the bounds of announced trust)
Participation above 95% Too easy — note the difficulty for the next cycle, keep the current event as is
Enhancement stone price beyond -10% in the 30 days post-event Strengthen sinks (limited shop), lower next quarter's inflation cap
Weekday playtime drops after the event Event-dependency signal — shore up weekday content appeal, adjust event frequency

The last row (weekday playtime dropping) is the signal most often missed. Watch only DAU (daily active users) during the event window and every event looks like a success. But if players do not come back on weekdays after the event ends, the event is feeding on the everyday appeal of the game. That is why §15.2.2's template carries "weekday playtime after the event" as a post_event_kpi slot from the very start. What is not measured cannot be prescribed for.


15.2.9 How Honest Can We Be About the Effects

An events chapter faces a strong temptation to drop in a table like "we ran a co-op event and retention rose from 30% to 50%." Numbers like that, unverified, erode the book's credibility. There are only three things this chapter can say.

First, direction can be stated from industry observation. Boosted attendance rewards lift short-term active user counts, co-op events strengthen community bonds, and limited packages lift revenue during the event window — that is the received wisdom of an industry that has watched live games for years. But how much swings widely with the game and its player mix, so carrying another company's numbers over wholesale is dangerous.

Second, an author's estimate is labeled an estimate. "Season pass completion target 70%," "fatigue builds past 10 days of event duration," "event production from half a day to one hour" — these are my experience-based estimates and unverified hypotheses. Do not memorize the absolute values; read the structure (template plus lint replacing blank-page design).

Third, only what is measurable gets promised as a KPI. Outcome metrics like retention are not driven by a single event, so I make no causal claims. What this workflow actually makes measurable is this — lint WARN counts (until reward violations reach 0), quarterly inflation totals (against the cap), repetition intervals for the same skeleton (in weeks), and per-event participation rates with post-event enhancement stone price movement. These four can be spoken in numbers at a meeting, not in "feelings."


15.2.10 Common Failures

Pattern Why it fails Prescription
Designing events from a blank page every week Slow, and results wobble Enter proven skeletons as template YAML (§15.2.2)
Wholesale delegation — "AI, make me a summer event" You get the average generic RPG event Fix the skeleton, vary only the slots (§15.2.3)
AI freely proposes reward amounts Inflation starts right there Lint enforces reward_policy (§15.2.4)
Reviewing candidates by eye only Quarter totals and repetition intervals get missed every time Automated checks with event_lint.py (§15.2.4)
Lint pass = straight to adoption Misses quarter rhythm and purpose conflicts The human gate watches the whole quarter (§15.2.5)
Trying to claw back rewards after the announcement Irreversible trust cost Every review before the announcement (§15.2.7)
Measuring only DAU during the event Misses the erosion of weekday appeal Post-event weekday playtime slot (§15.2.8)

The fifth is the one most often missed. Send a lint PASS straight to announcement and you lose the place to shake a candidate like Flood Leveelegal, but one that drains the quarter-end reward headroom to 0. Code sees the legality of one event; humans see the rhythm of the whole quarter.


15.2.11 Try It Yourself — One Step You Can Take Today

If You Work Solo, Do Just This Much: You do not need the lint code. Pick one event skeleton that shows up often in your own game (or in a live game you love) and write a template YAML in the §15.2.2 format by hand (the three fields that matter are core_loop, slots, and reward_policy). Then attach the §15.2.3 prompt, pull 5 variation candidates, pick the one whose rewards feel too generous, and push back: "this blows past this month's reward budget — lower it and try again." What adoption and rejection actually bundle together sinks in through your hands.

If you are on a team, start with this one step. Enter the 3–4 event skeletons you run most often as template YAML, and build the three lines of event_lint.py — reward cap, quarterly inflation total, repetition interval — in code first. With just the templates and those three lines, you head off the two most common failures: blank-page design every week, and rewards set by gut. This workflow is the first hands-on implementation of §15.1.5's three progressive-application skeleton elements — an event template and season rule library, an AI event candidate generator, and automated post-event measurement.


Key Takeaways

Next Chapter Preview

15.3 100 Feedback Items into Topics — Clustering Goes to the LLM, Priorities Stay with People

Primary audience: designers and directors responsible for user-facing live operations (live ops) on a mid-sized team (10–50 people) Scaled-down version for solo/hobbyist readers: §15.3.7 "If You're Solo, Just This Much"

Let me be honest up front. I have not spent long stretches — a year or two at a time — personally owning post-launch live ops. Much of this chapter is industry observation and adjacent experience layered on top of 24 years in the field. So this chapter does not presume to declare "this is how you run live ops." Instead, I take the input → AI → verification → human decision cycle that was proven in pre-launch content production, plug in a new input — user feedback — and run it through one full cycle to see what comes out. The skeleton of the tooling is the same as the city_hunting_generator in §6.2; only the input changes, from "city metadata" to "100 items of user feedback."

The first week of live operations looks much the same everywhere. Forum posts, Discord messages, CS (customer support) tickets, and store reviews pile up by the hundreds to thousands a day. No human can read them all, and if no one does, the same bug report gets buried fifty times over. This chapter covers a method where an LLM clusters that pile into topics and scores the sentiment, so that people step in only for the priority decision: "so what do we fix this week?"


15.3.1 Feedback Is "Classification Input," Not "Reading Material"

Every live ops textbook has the table that splits feedback into four channels (in-game surveys, forums/Discord, store reviews, CS tickets) and four types (bugs, requests, complaints, praise). It is all true. The problem is that memorizing the table gives you no answer to "how do we handle the 412 items that came in today?" As long as you treat feedback as something for humans to read and classify, feedback volume always beats ops headcount.

Shift the perspective. One piece of feedback is structured input: a record with five slots, {source, raw text, topic, sentiment, severity}. Seen this way, the nature of the work changes. It is not "read everything" but "cluster by topic and rank by priority." Topic clustering and sentiment scoring are tedious for humans, and our criteria drift every time we do them — a machine applies the same yardstick identically to all 100 items. This is exactly the kind of work an LLM does better than people. The division of labor that mass-produced 30 cities in §6.2 (rulebook = deterministic, body text = AI, review = human) holds here unchanged. Only one thing differs: the human's job at the end is not "review the text" but "decide the priorities."

One note on the distribution of feedback types. Users who write voluntarily skew toward the dissatisfied, not the satisfied. Happy customers leave quietly; unhappy customers come back to the counter. So the sentiment distribution on forums and reviews tends to lean more negative than the actual satisfaction of your whole player base (my observation — the exact size of the bias varies by game, channel, and period, so read it as a direction, not an absolute number). Keep this bias in mind, and when you see "60% negative" in the clustering results you won't misread it as the game going under.


15.3.2 [Worked Transcript] 100 Feedback Items → Topic Clusters + Sentiment

Let's actually run one cycle end to end. The input is 100 feedback items collected from four channels over one week; the output is topic clusters, sentiment, and priorities. The input prompt can be copied as is, and the output below is a reconstruction of the format from an actual classification session.

Step 1 — Input: Turn Feedback into a Machine-Readable Table

Normalize the raw text scraped from each channel into one record per line. This is not new writing — it is extraction and cleanup only.

{"id": "fb_0001", "src": "discord",     "text": "Failed 50 times going for +12 enhancement. Are these rates even right? I want a refund"}
{"id": "fb_0002", "src": "store_review","text": "Graphics are pretty but the lag is so bad I get disconnected every guild war"}
{"id": "fb_0003", "src": "cs_ticket",   "text": "I paid but the diamonds never arrived, order number attached"}
{"id": "fb_0004", "src": "forum",       "text": "When is the new archer class coming T_T you promised it at pre-registration"}
{"id": "fb_0005", "src": "discord",     "text": "First week since launch and the ops team communicates well, notices are fast. Keep it up"}
{"id": "fb_0006", "src": "store_review","text": "One boss (Heukrang) does absurd damage. Full gear and it one-shots me. Balance patch please"}
{"id": "fb_0007", "src": "cs_ticket",   "text": "Can't progress at tutorial step 5, the button won't press (device: Galaxy A series)"}
// ... fb_0008 ~ fb_0100 (omitted)

At the input stage the records leave topic·sentiment·severity empty. Filling those blanks is the LLM's job in step 2.

Step 2 — The Prompt: Demand Clustering, but Force the Labels, the Format, and the Escape Hatch

Cluster the attached feedback_100.jsonl (100 feedback items from one week) by topic, and score sentiment as well.
Pick topics only from this list (no free generation): Enhancement/Rates, Balance, Server/Performance, Payments/Refunds,
New Content Requests, Tutorial/Onboarding, UI/Controls, Praise/Support, Other. If 'Other' exceeds 8 items, also propose new topic candidates.
Sentiment is Negative·Neutral·Positive; severity is S1·S2·S3·S4.
// (Intent: S1 only when specific, reproducible, and feature-blocking. A merely strong complaint is S2)
If you're not confident about an item, leave it in 'Other', mark a ★ at the end of its id, and pass it to me — don't force-fit it.
Give me just two tables — Table A (per item): id·topic·sentiment·severity / Table B (per topic): topic·count·negative %·1 representative quote·has S1.

What determines the classification quality in this prompt is neither the topic list nor the sentiment scale, but three control devices: pinning the labels, nailing down the format, and leaving open an escape hatch. Let the LLM free-generate topics and it invents mushy labels you cannot prioritize with, like "general dissatisfaction" or "requests to improve the game." The allowed-label list (rule 1), the severity definitions (rule 3 — the narrow definition of S1 is the key), the output format (two tables), and "if unsure, mark ★ and hand it to a human" (rule 5). Drop any of these four and the table degenerates from a priority tool into yet another pile of reading material. This label list is itself the category scheme of the feedback L4 layer from §15.1.5 — labels must stay fixed so that each quarter's trends can be compared along the same axis.

Step 3 — Claude's Output (Reconstructed, Excerpt)

[Table A] Per-item classification (excerpt)
id       | Topic                | Sentiment | Severity
fb_0001  | Enhancement/Rates    | Negative  | S2
fb_0002  | Server/Performance   | Negative  | S2
fb_0003  | Payments/Refunds     | Negative  | S1
fb_0004  | New Content Requests | Neutral   | S3
fb_0005  | Praise/Support       | Positive  | S4
fb_0006  | Balance              | Negative  | S1
fb_0007  | Tutorial/Onboarding  | Negative  | S1
fb_0042★ | Other                | Neutral   | S3
...

[Table B] Per-topic aggregation (all 100 items)
Topic                | Count | Neg % | Representative quote                       | Has S1
Server/Performance   | 24    | 92%   | "disconnected every guild war"             | Yes (partially reproduced)
Enhancement/Rates    | 21    | 95%   | "failed 50 times going for +12"            | No
Balance              | 14    | 86%   | "Heukrang one-shots full gear"             | Yes
New Content Requests | 12    | 8%    | "when is the archer coming"                | No
Tutorial/Onboarding  | 9     | 78%   | "step 5 button won't press (Galaxy A)"     | Yes
Payments/Refunds     | 8     | 100%  | "the diamonds never arrived"               | Yes
UI/Controls          | 6     | 67%   | "skill buttons too small, mis-taps"        | No
Praise/Support       | 4     | 0%    | "the ops team communicates well"           | No
Other                | 2     | -     | (fb_0042★, fb_0067★ awaiting human call)   | -

Overall sentiment: Negative 71 / Neutral 18 / Positive 11
'Other' at 2 items (under 8) → no new topic proposal needed
Low-confidence ★ marks, 2 items: fb_0042, fb_0067 (human judgment requested)

The most valuable part of this output is not the tables but the two lines at the bottom — the "2 items marked ★." That is where the LLM reported what it could not cluster and handed it to a human. It is the same design as in §6.2, where the AI flagged the NPC "Grem" as ambiguous on its own. A good prompt makes it possible for the AI to say "I am not confident about this."

Step 4 — Verification and Veto (the Human's Seat)

You must not accept this output as is. One spot actually snagged.

All 21 items in the Enhancement/Rates topic were classified S2 (complaint). But one of them, fb_0001, carries "I want a refund." The LLM read this only as "strong complaint (S2)." Here the human steps in. A complaint about enhancement rates — as long as the data shows the rates working exactly as specified — is not an S1 incident. Dissatisfaction with odds that behave as specced is a design and perception problem, not a bug. The LLM's S2 call is correct. But the "refund demand" signal needs to be cross-linked to the payments topic so CS can handle it separately. The LLM assigned a single topic label per item and missed the case where one item straddles two topics.

So I re-request.

Added rule: when one item straddles two topics (e.g., enhancement complaint + refund demand), write the
secondary topic in a 'cross' column in addition to the main topic. Re-output Table A with a cross column added.
However, keep enhancement-rate complaints themselves at S2, not S1, as long as the data shows the rates match the spec.

One round trip and it is done. The LLM re-answered fb_0001 as topic=Enhancement/Rates, cross=Payments/Refunds, severity=S2, and I read the two ★ items myself, reassigning fb_0042 to UI/Controls and fb_0067 to Tutorial/Onboarding. Reading and classifying 100 items from scratch by hand takes half a day; an LLM draft + human review + one round trip stays under an hour (my estimate, an unverified hypothesis — the exact savings depend on feedback volume and channel count, so read it not as absolute times but as the structural difference between "by hand from scratch" and "draft + review").


15.3.3 Priorities Are Not the LLM's to Give — The Human's Seat

Here I draw the decisive line. Table B above only says "which topic has how many items, and how negative they are." "So what do we fix first this week?" is something the LLM cannot give you. That decision entangles cost, schedule, and the game's vision, and the responsibility for it rests with the director.

Two ops teams can look at the same table and reach opposite decisions. By count alone, Server/Performance (24 items) and Enhancement/Rates (21 items) rank first and second. But priority does not follow count order. The reason is severity and reversibility.

flowchart TB
    A["100 feedback items
(4-channel normalized jsonl)"] --> B["LLM clustering
scores topic, sentiment, severity"] B --> C{"Human review
★ items, cross, misclassification"} C -->|Re-request| B C -->|Confirmed| D["Per-topic aggregate table
count, negative %, has S1"] D --> E["Priority decision
(director — not the LLM)"] E --> F1["S1 incidents: immediate hotfix
payment and tutorial blockers"] E --> F2["S2 complaints: check data,
then design judgment"] E --> F3["S3 requests: quarterly backlog,
codified into voice slots"] F1 --> G["Reply cycle
(§15.3.4)"] F2 --> G F3 --> G classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; class A,D data; class B ai; class C,E human;

In this flow, human hands touch only two places: the review gate in the middle (judging ★ items, cross links, and misclassifications) and the priority decision at the bottom. The tedious classification of 100 items in between is the LLM's run. And the actual logic of the priority decision is not the count — it is the following three axes.

Topic Count Priority call (the director's seat)
Payments/Refunds (S1) 8 First priority. Few items, but feature-blocking + irreversible (money). 24h hotfix
Tutorial/Onboarding (S1) 9 Second. Directly tied to new-user churn. Reproduced on specific devices → patch
Server/Performance 24 Third. Highest count, but infrastructure work = long schedule. No hotfix possible; next week
Enhancement/Rates (S2) 21 Hold. If the data matches the spec, not a bug. Reviewed separately as a design decision
New Content Requests 12 Backlog. 8% negative (= positive anticipation). Codified into the quarterly voice slot

Server/Performance, first by count, drops to third priority because it is infrastructure work a hotfix cannot cover; Payments/Refunds, sixth by count, rises to first because it is an irreversible incident with money on the line. The LLM cannot do this reordering. The LLM delivers only the fact: "payments, 8 items, 100% negative." The decision that this is priority one belongs to a person who knows the costs, the legal risks, and the game's vision. This is what §15.1.5's "AI produces the classifications and candidates; people focus on adoption and vision decisions" actually looks like in the feedback domain.


15.3.4 Replies — An Irreversible Step, so the Review Gate Weighs Heavier

Once priorities are set, you reply to users. In live ops, the absence of replies is where trust takes its biggest damage. Even when there is nothing to say, "under review" beats silence. Reply drafts, too, can be pulled from the LLM per topic.

[Reply drafts — LLM output, per topic]

Here is the one decisive difference from §6.2: sending a reply is an irreversible step. A city NPC can be discarded and regenerated, but a notice or reply text a user has seen once cannot be taken back. If you auto-send "granted within 24 hours" and it actually takes three days, that promise remains in the community as an irreversible mark. So the irreversible-step principle of §15.1.4 operates more heavily in the feedback domain than in any other. The LLM produces the auto-reply drafts, but not a single character goes out automatically before passing the CS review gate. The reviewer checks only two things: whether the schedule promises (24h, next week) match the actual work schedule, and whether sensitive cases (legal disputes, refund disputes) have slipped into the auto-send pool. It is a seat for the judgment that lint cannot catch.

Step Reversibility Who
Feedback clustering and sentiment scoring Reversible (free to rerun) LLM
Topic review and priority decision Reversible (until confirmed) Human (director)
Reply draft generation Reversible (discard and rewrite) LLM
Reply send / notice posting Irreversible (seen by users) Human (after CS review)

15.3.5 Codify User Voice into the Quarterly Retrospective

To keep the same feedback from swinging to different decisions every quarter, the clustering results must be codified as a fixed input slot in the quarterly retrospective. Not an offhand "lots of enhancement complaints lately," but a table aggregated along the same label axis every quarter, sitting inside the retrospective table. This is where the decision in §15.3.2 — banning free-form labels and fixing an allowed list — pays off.

2026 Q2 user voice (LLM auto-aggregated, quarterly cumulative)

Roughly 5,000 items clustered across 4 channels (counts are actual quarterly tallies — not embellished)

Top negative topics:   Enhancement/Rates > Server/Performance > Balance > Payments/Refunds
Top requested topics:  New class > New hunting grounds > Guild system > UI improvements
Quarterly sentiment:   Q1 negative 68% → Q2 negative 71% (slightly worse — driven by the enhancement topic)

This table becomes the input to the quarterly decision. The decision itself belongs to the director; the input belongs to the users. Read the quarterly trend ("Q1 68% → Q2 71%") as direction only. The signal is not any single quarter's absolute value but the direction of change along the same label axis. If the negative % rose, trace back which topic pulled it up and connect that to next quarter's priorities. The draft of this quarterly report is itself something the LLM writes in natural language, with the human adding only the decision comments — the actual seat of the auto-drafted quarterly report described in §15.1.5.


15.3.6 How to Handle Numbers Honestly

A live ops chapter carries a strong temptation to insert a table like "after adopting the feedback cycle, NPS (Net Promoter Score) rose from 20 to 45." I have never measured that causation, so I do not write it. This book's principle is one of three.

First, actual tallies are written as is. The per-topic counts in §15.3.2 (server 24, enhancement 21, payments 8) and the quarterly cumulative in §15.3.5 are values counted item by item from classification results, not ratios groomed to look good.

Second, estimates are labeled as estimates. "Classifying 100 items: half a day → one hour" (§15.3.2) and "forum sentiment skews negative" (§15.3.1) are my experience- and observation-based estimates, unverified hypotheses. Do not memorize the absolute values; read the direction (feedback volume always beats headcount; voluntary posts lean toward complaints).

Third, only what is measurable gets promised as a metric. What a feedback cycle can actually measure is not outcome satisfaction (NPS) but process metrics — the backlog of unclassified feedback (target: 0), the lead time from S1 discovery to hotfix, reply response time, and the share of the 'Other' topic (when the allowed labels fail to capture reality, 'Other' balloons). These four let you speak in a meeting with numbers, not "feelings."


15.3.7 Try It Yourself — One Step You Can Take Today

If you're solo, just this much: You need no CS system and no dataset. Copy 20–30 store reviews or community posts for your game (or a game you love) by hand into a jsonl ({"id":..., "src":..., "text":...}), paste the prompt from §15.3.2 as is, and run it once. In the resulting Table B, find one case where the "topic with the highest count" and "the topic you want to fix first" differ, and write one line on why they differ — and you will feel in your bones why priorities are a human's job, not the LLM's.

If you are on a team, start with this one step. First, fix the extraction script that collects 4-channel feedback into one-record-per-line jsonl, and the allowed topic label list from §15.3.2. Only with fixed labels can LLM classification and human classification measure along the same axis, and quarterly trends be compared. Auto-replies come after that — replies are irreversible, so never wire them to auto-send without a CS review gate.


Key Takeaways

Next Chapter Preview

Part 16 · Communicator

16.1 Combat TF Operations — Only Decisions Leave the Isolated Workspace as Canon

Thursday, 4 p.m. The combat TF meeting ended and the seven of us scattered back to our desks. The whiteboard still holds the traces of whether to cut the global cooldown from 0.8 seconds to 0.5. The senior balance designer said, "My sim says 0.5 is right." The code lead said, "At 0.5 the server tick can't keep up." The UI designer said, "I can't speak to either of those, but the cooldown gauge gets too narrow."

All three of them are right. And if all three start writing their own conclusions into their own discipline's documents, those three documents will contradict each other by next week. The balance sheet says 0.5, the code spec says 0.8, the UI guide says 0.6. Nobody can tell which one is the canon.

The combat TF exists precisely to absorb this collision in one place. And only the product of that absorption — a single decision — should rise into the canonical docs, the documents that count as the single source of truth. The debris of everything else debated has to end inside an isolated workspace. This chapter covers that mechanism of isolation and absorption.


16.1.1 A TF Is Not a Permanent Department but an Isolated Workspace

A major combat system overhaul never ends within one discipline. Touch the global cooldown alone and balance (numbers), code (server tick), UI (gauge rendering), animation (motion length), and sound (hit feel) all shake at once. Run an agenda item like this through each discipline separately and the decision stretches out by two to four weeks — and even once decided, the disciplines end up misaligned.

A TF (task force) is a unit that temporarily gathers several disciplines into one workspace to prevent that misalignment. The key words are "temporarily" and "isolated." If you let the TF's debate flow straight into the company's canonical document system, unverified discussion, rejected proposals, and in-experiment numbers contaminate the canon. So we create an isolated workspace in SVN under a number that starts with 95_.

95_BattleTF. The 95 range is our convention for short-lived TF workspaces. Regular canonical docs use the 10s and 20s; the 90s signal "temporary, isolated, scheduled to end." The folder number alone instantly communicates: this is not canon — do not cite numbers you saw here.

The rules of isolation are simple.

This is why a TF that hardens into a permanent department is dangerous. Once the isolation breaks, unverified numbers from the TF workspace start getting cited as if they were canon, and the same decision gets broken again every quarter in a different room.


16.1.2 From Isolation to Absorption: The Full Flow

One cycle of the combat TF opens an isolated space, accumulates debate, experiments, and decisions inside it, and at close-out absorbs only the decisions into the canon.

flowchart TD
    A["Combat agenda item arrives
(global cooldown 0.8→0.5?)"] --> B["Open 95_BattleTF isolated space
(SVN 95_ number range)"] B --> C["Artifacts accumulate inside isolation
minutes, experiment sheets, rejected proposals, raw notes"] C --> D{"Decision reached?"} D -->|Not yet| C D -->|Confirmed| E["Update TF_결정사항_요약.md
(inside the isolated space)"] E --> F{"TF ends?"} F -->|Continues| C F -->|Ends| G["TF_결정사항_요약.md
this one file alone promoted to canonical docs"] G --> H["Everything else
demoted to 95_BattleTF/archive/"] G --> I["Art team gets html only
(md source not shared)"] H --> J["Isolated space closed"] style B fill:#fff3cd,stroke:#d39e00 style G fill:#d4edda,stroke:#28a745 style H fill:#f8d7da,stroke:#dc3545

The agenda item enters at the top left, all the noise gets processed inside the yellow isolated space, and only the green box — the decision summary — exits into the canon. Red means demotion. This one diagram is the whole of how a 95-range workspace operates.


16.1.3 The Worked Transcript — Making the Decision Summary Absorbable

The most labor-intensive job at TF close-out is sifting a quarter's worth of minutes and experiment sheets down to "only the decisions that go to canon." The debate runs long, rejected and confirmed proposals sit mixed together, and the same number is written slightly differently in every meeting. Done by hand, the close-out alone eats a full day.

Below is the prompt I actually ran, Claude's raw output, and the full sequence of how I verified, rejected, and re-requested. I print it as is, without summarizing.

First Prompt (Full Text)

From the 6 sets of 95_BattleTF meeting minutes below, pull only the confirmed decisions to promote to canon
and draft TF_결정사항_요약.md. The TF is closing soon.
Confirmed items only (drop rejected, in-experiment, and "let's revisit later"); format each decision as
decision ID, topic, confirmed value, evidence (data source), decider, confirmation date.
If the same topic changed values across meetings, keep only the final value and drop the earlier ones.
If confirmation is ambiguous, don't drop it — set it aside as [Needs Review]; if no evidence source is given, mark it [No Evidence on Record].

Minutes:
--- Minutes 2026-W14-mon ---
Item 1: Review global cooldown 0.8→0.5. Per balance sim (sim_gcd_0407.xlsx),
0.5 gives DPS variance -12%. Code lead: requests holding 0.5 due to the server tick (33ms) limit.
Decision: first-pass agreement on 0.6 (provisional). Code side to load-test 0.6 by next meeting.
Item 2: Dodge i-frames 12f→10f. Anim director: motion is 8f, so 10f breaks visual continuity.
Decision: keep 12f, confirmed. Decider: design lead.
... (nothing elided here; in practice all 6 sets of minutes went in verbatim) ...
--- Minutes 2026-W16-thu ---
Item 1: Global cooldown. 0.6 load test passed (18% server tick headroom).
Decision: 0.6 confirmed. Decider: design lead. Evidence: sim_gcd_0407.xlsx + load_0415.log.
Item 3: Kill reward coins ×1.2. Balance proposal, no evidence. On hold.

Claude's Raw Output (As Is)

# TF_결정사항_요약.md (draft)

## Confirmed Decisions

### D-GCD-01 | Global Cooldown
- Confirmed value: 0.6s (was 0.8s)
- Evidence: sim_gcd_0407.xlsx (DPS variance impact), load_0415.log (18% server tick headroom)
- Decider: design lead
- Confirmed on: 2026-W16-thu
- Note: 0.5 was proposed in W14 but converged to 0.6 due to the server tick limit. The 0.5 proposal is discarded.

### D-IFRAME-01 | Dodge i-frames
- Confirmed value: keep 12 frames
- Evidence: [No Evidence on Record] — only a qualitative judgment is recorded (8f anim motion, visual discontinuity)
- Decider: design lead
- Confirmed on: 2026-W14-mon

## [Needs Review]

### R-COIN-01 | Kill Reward Coins ×1.2
- Status: on hold (balance proposal, no supporting data)
- Judgment: cannot confirm without evidence. Recommend transferring to the balance team as a standalone item after the TF closes.

My Verification — What I Accepted and What I Rejected

I went through the three items one by one.

D-GCD-01 passed as is. Dropping the provisional 0.5/0.6 from W14 and keeping only the confirmed 0.6 from W16 is exactly right. Rule 3 was followed properly. Bundling the two evidence sources together is also correct.

D-IFRAME-01 I rejected. Claude classified it as "confirmed," but its evidence reads [No Evidence on Record]. If a decision enters the confirmed list on a qualitative judgment alone (visual discontinuity), anyone reading the canon later cannot find the grounds for "why 12f." This is a collision case between rules 1 and 5 — Claude saw the decider had confirmed it and called it "confirmed," but I had to apply our docs policy: a confirmation without evidence does not go to canon. That policy is written nowhere in the minutes, so Claude had no way to know it.

R-COIN-01 was classified correctly, but the prescription overreached. "Recommend transferring to the balance team as a standalone item" is a procedure Claude invented. Our company has no such transfer track. I accepted the classification (needs review) and threw out the prescription sentence.

Second Prompt (Re-Request)

Redo it with my verification applied. Leave D-GCD-01 as is,
and since D-IFRAME-01 has only a qualitative judgment for evidence, pull it out of confirmed and move it to [Promote After Evidence Added].
For R-COIN-01 keep the classification but drop prescription sentences like "recommend transferring" — you don't know our company's procedures, so record only the status and the facts.
And don't invent procedures, tracks, or team names that aren't in the minutes; if you don't know, leave it as [Design Lead Decision Needed].
// (intent: promotion condition for D-IFRAME-01 = motion frame measurement data from the anim director attached to the minutes)

Second Output (Final)

# TF_결정사항_요약.md

## Confirmed Decisions (To Be Promoted to Canon)

### D-GCD-01 | Global Cooldown
- Confirmed value: 0.6s (was 0.8s)
- Evidence: sim_gcd_0407.xlsx, load_0415.log (18% server tick headroom)
- Decider: design lead / Confirmed on: 2026-W16-thu
- Note: the 0.5 proposal was discarded due to the server tick limit.

## [Promote After Evidence Added]

### D-IFRAME-01 | Keep Dodge i-frames at 12f
- Only a qualitative judgment is recorded (visual discontinuity against the 8f anim motion).
- Promotion condition: motion frame measurement data attached to the minutes.

## [Design Lead Decision Needed]

### R-COIN-01 | Kill Reward Coins ×1.2
- Status: on hold. Balance proposal, no supporting data. No confirmed value.

From this final version, only the "Confirmed Decisions" section — the single item D-GCD-01 — was promoted to the canonical docs. The other two sections went down to archive/. A quarter's worth of six meeting minutes was absorbed into one line of canon.

This is where what the AI did and what it could not do split. The AI tracked value changes on the same topic across six sets of minutes, separated out the rejected proposals, and flagged missing evidence — and that line-by-line cross-checking of six documents is exactly the kind of simple repetition that slips through human hands. But applying the policy that a confirmation without evidence does not go to canon, knowing the company fact that no transfer track exists, and making the final confirmed-or-held call — all of that was human. Delete the AI's paragraphs and what disappears with them is the extraction and sorting labor; the decision of what goes into the canon stays in human hands either way.


16.1.4 External Requests Come In Through 3-Track Triage

Not every agenda item entering the TF originates inside it. Requests come in from the publisher, art outsourcing vendors, and the business team: "It's combat-related — please handle this." Accept these indiscriminately as TF agenda items and the TF turns into an external complaints desk.

So external requests are triaged into three tracks the moment they arrive. Only those requiring a combat decision go into 95_BattleTF; work that ends within a single discipline is handled solo by its owner; out-of-scope or under-evidenced requests get a written reason and are returned or held. Only the first track enters the TF — that is the first line of defense against the TF degenerating into a complaints desk. The triage itself is a human judgment, but having the AI do a first pass over the incoming request text and tag "how many disciplines does this touch" is perfectly fine.

The ruling order, the worked transcript, and the per-track follow-up of this three-way triage (request-triangulate) are the subject of the next chapter, 16.2. Here I only pin down the entrance rule: the TF accepts the first track only.


16.1.5 The Art Team Gets html Only — Zero md to Learn

Once a TF decision is promoted to canon, it gets shared with the relevant teams. There is one asymmetry here. The art team does not receive the Markdown source (.md); they receive only the rendered html.

The reason is simple. The art team only needs the outcome of the decision. "Re-fit the cooldown gauge width to the 0.6-second baseline" — that one line is everything they need. The md source contains the decision ID scheme, atom references, traces of the rejected 0.5 proposal, and evidence data file names. That is a working language shared by design and code, not something art should have to learn.

Hand over the raw md and the art team pays two costs. First, they spend time deciphering a notation system irrelevant to them. Second, they can mistake unverified or rejected information for a decision. html blocks both — only the cleanly rendered decision outcome is visible, and the internal notation is filtered out during the build.

Written as a principle: the working language (md) circulates only within the disciplines that speak it; only the deliverable (html) goes outside. It is the same philosophy as the TF workspace isolation (the 95 range). Keep the raw material inside, and send out only the absorbed result.


16.1.6 The Foundation of TF Operations — Five Principles

For the isolation-and-absorption mechanism to run, five operating principles have to sit underneath it. Drop any one of them and the TF collapses into a debating chamber.

When the five principles operate together, the isolated 95-range space becomes a decision factory instead of a debating chamber.


16.1.7 Common Pitfalls

Here are the pitfalls that recur from the mid-stage of TF operations onward, with their remedies.

Pitfall Symptom Remedy
Degenerates into a discussion club Opinions exchanged, no decisions Force N decision slots per meeting
Authority encroachment TF intervenes in other disciplines' decisions Clarify the decision rights table
Member overload Overlapping membership in 5–6 TFs erodes day jobs Cap total TF participation at 8 hours per week
Becoming permanent Same meetings repeat with no dissolution Quarterly re-evaluation
Isolation leak Unverified 95-range numbers cited as if canon Promote only the one decision summary to canon
External disconnect Decisions not shared outward Promote to canon + deliver html

The isolation leak is the quietest and the most dangerous. When the folder number convention collapses, everything collapses.


16.1.8 Measurement — What the TF Absorbs

From my Project A operating records I carry over only directions and ratios. The figures below are not absolute values but directions of change when running a TF versus not having one — absolute cycle times vary with team size and build cadence (observations from my own environment).

Item Without TF With TF Direction
Cycle for one combat decision Per discipline, separately; weeks A matter of days Shorter
Cross-discipline conflicts after a decision Many per quarter Few per quarter Down
Game director escalations Several per week 1–2 per week Down
Cross-discipline information sharing Sporadic Fixed via minutes and canon promotion Systematized

What gets reclaimed most is the game director's time. Because the TF absorbs cross-discipline decisions inside the isolated space, fewer conflicts climb all the way up to the director's desk. A TF is, in the end, a device that takes the cross-discipline consensus-building the director used to mediate case by case and pulls it down into a single workspace.


Key Takeaways


Beyond Games. The principle — from an isolated workspace, only decisions get absorbed into the canon — applies as is to any cross-department project that has nothing to do with games. Picture a TF where marketing, legal, and sales work together on revising the terms of service. Keep the minutes, review comments, and rejected draft clauses in a temporary folder on the shared drive (an isolated space like 95_약관TF, a terms-revision TF folder), and when the TF ends, move only the single final confirmed-language file (최종_확정문구.docx) into the company's canonical document library and send everything else down to the archive. Six months later, when someone asks "why did we settle this clause this way," this is what prevents an unconfirmed draft from passing itself off as the canon.


Try It Yourself — Absorbing Decisions at Quarter Close

setup - In SVN (or any folder system), create the isolated space 95_BattleTF/ and gather a quarter's worth of meeting minutes inside it. - Create 95_BattleTF/archive/ in advance (where the demoted material will go).

prompt - Paste this chapter's first prompt together with the full minutes. Core rules: ① confirmed decisions only ② only the final value per topic ③ if ambiguous, don't discard — set it apart with a label ④ flag missing evidence ⑤ do not invent company procedures or team names.

verify - Go through the output's "confirmed" classifications one by one. Pull any item whose only evidence is a qualitative judgment back out of "confirmed" (apply your canon-promotion policy). - Check the AI's prescription sentences (transfers, tracks, recommendations) for procedures that do not actually exist, and delete them. - Copy only the "Confirmed Decisions" section into your canonical docs, and send the rest down to archive/.


16.1.9 Solo Scale-Down

Isolation and absorption hold just as well for a solo developer working alone. Just replace "TF" with "the several roles inside my head."

With an isolated space, folder location alone tells you whether a number is confirmed or still an experiment. Even working alone, this is the cheapest way to avoid handing the same confusion down to your future self.

16.2 Collaborating with Other Disciplines — Sorting External Requests into 3 Tracks

On a Tuesday morning, the team messenger pinged three times almost at once.

Art lead: "The combat effect colors feel really drab right now — can we go brighter?"

QA lead: "There's a case where the guild attendance reward gets paid out twice. Repro video attached."

Publisher contact: "Please apply the Islamic-market cultural guidelines to the Southeast Asia build. Before next quarter's review."

The three messages were about the same length. But one was a 30-minute job, one was an incident that needed the code lead pulled in immediately, and one was an external schedule item that had to be slotted into the quarterly plan. Treat them with the same weight just because they landed in the same inbox, and you spend half a day on the 30-minute job while the actual incident sits untouched until evening.

The requests that reach a game designer differ in grain as much as the disciplines they come from. The problem is that they all arrive in the same shape: a one-line message. This chapter covers the work of splitting those one-liners into three tracks the moment they arrive. The moment the track is decided, so is what you stop right now and what you push to later.


16.2.1 Collaboration Decides the Core Work

A game designer doesn't write the code, paint the art, or compose the sound. We write specs, convey intent, and verify results. Every deliverable comes out through another discipline's hands. So the quality of collaboration directly determines the quality of design output.

On Project A — the mobile-first MMORPG I direct, with a mid-sized team (10–50 people) — the disciplines a designer collaborates with day to day look like this.

Game designer 40–60% of time Dev (code/tools)Daily Art2–3× a week Sound1–2× a week Animation1–2× a week QAWeekly + MS Live ops & CSWeekly Publisher/platform1–2× a quarter

Seven disciplines, meshing at rhythms from daily to quarterly. 40–60% of a designer's desk time goes into this collaboration. The core work — design itself — gets roughly the other half. Which means cutting collaboration time is the same thing as growing core-work time. And the biggest drain on collaboration time is failing to classify incoming requests and pouring energy into the wrong place.


16.2.2 The Three Grains Hidden in a One-Line Request

Back to the three messages. On the surface they all read "please do X." Underneath, three different natures are hiding.

I call each of these three grains by one word: align, defect, schedule. The work of pushing every incoming request into one of these three first is something we have codified on Project A as a workflow named request-triangulate. The name comes from triangulation: you fix the position of one point (the request) by surrounding it with three reference points (the nature of the discipline, urgency, external dependency).

The classification flow looks like this.

flowchart TD
    A[1 external request arrives] --> B{Tied to an external
deadline or contract?} B -- Yes --> S[Track-S: schedule] B -- No --> C{A user-impacting
bug or defect?} C -- Yes --> D[Track-D: defect] C -- No --> E{A taste/intent matter
to settle by agreement?} E -- Yes --> F[Track-A: align] E -- No --> G[On hold:
request more info] S --> S1[Fold into quarterly roadmap
secure ample lead time] D --> D1[Rate priority P0~P2
loop in code lead immediately] F --> F1[Convey intent only
delegate expression to that discipline] style S fill:#fde2c4,stroke:#c98a3a style D fill:#f6c6c6,stroke:#c25151 style F fill:#c9e4d0,stroke:#4f9d6a style G fill:#e0e0e0,stroke:#888

The order of the questions is the point. Schedule dependency is asked first because, for anything tied to an external deadline, lead time outranks internal judgment. Classify a task whose quarterly review is 3 weeks away as "we'll align on it later," and by the time alignment ends the deadline is on top of you. Defects come second because anything already affecting users always outranks taste discussions. Alignment comes last. Only when something is not urgent, not bound to an external party, and not hurting users does "let's take our time and align" become legitimate.

If all three questions come back "no," that is not a classification failure — it is missing information. Don't force a track; put it on hold and ask back. One sentence like "Does this have to make the next build, or is a review enough for now?" usually settles the track.


16.2.3 Worked Transcript: Splitting the Inbox into Tracks

Do this classification only in your head and it collapses on a busy day. So I run it differently: throw the whole batch of incoming requests at the AI for a first-pass classification, then verify and correct the result. Below is one actual run, reproduced almost verbatim. No summarizing — the parts the AI got wrong and the parts I rejected are left in.

The Full Prompt

Classify the external requests below into three tracks — Track-S (schedule: tied to external deadlines, contracts, reviews),
Track-D (defect: user-impacting bugs, rated P0~P2), Track-A (align: taste and intent, closed by agreement).
Check external dependency → user impact → taste/intent in that order and lock in the first track that matches;
if none match, put it on hold and give me 1 question to ask back.
Output per request: | Track | Rationale (1 line) | Next action (1 line) |.

[Request list]
1. Art lead: combat effect colors look drab. OK to go brighter?
2. QA lead: guild attendance reward paid out twice. Repro video available.
3. Publisher: apply cultural guidelines to the Southeast Asia build. Before next quarter's review.
4. Sound director: boss BGM intro lags combat start timing by 0.5 seconds.
5. Live ops team: many user reports saying "the guild chat font is too small."

Claude's Raw Output

| # | Track | Rationale | Next action |
|---|------|------|-----------|
| 1 | Track-A | Color tone is taste/intent territory; no external or bug angle | Convey only the intent (combat tension); delegate saturation to art |
| 2 | Track-D | Duplicate rewards = direct impact on user assets, reproducible | Rate P0, loop in code lead immediately, consider hotfix |
| 3 | Track-S | Tied to an external deadline: the quarterly review | Fold into quarterly roadmap; break guidelines into items, then estimate lead time |
| 4 | Track-A | BGM timing looks like a matter of directing taste | Align with the sound director |
| 5 | Track-D | Font readability is a UI defect | File as a P2 bug |

Human Verification / Rejection

I stopped at #4 and #5.

#4 rejected. The AI read the BGM intro lag as directing taste (Track-A). But "lags combat start timing by 0.5 seconds" is not taste — it is a synchronization defect. It is most likely not the audio asset itself; the trigger point on the Behavior Tree (BT) and the audio cue have probably drifted apart. That can't be closed by agreement — it means touching timing data on the animation and code side. → Corrected to Track-D. P2, though, since no user assets are affected.

#5 demoted to hold. The AI jumped straight to UI defect (Track-D) on "the font is too small." But whether that is a defect or taste can't be settled from the message alone. If the font renders exactly to design spec and merely feels small, that is closer to align (Track-A); if it is rendering broken, smaller than spec, that is a defect (Track-D). → On hold. Asked live ops back: "Does it actually render smaller than the spec font size, or is this a request to raise the spec itself?"

Re-request

The added instruction in the re-thrown prompt, reflecting the two rejections, was short.

Reclassify #4 — '0.5s lag behind combat start' — as a synchronization defect (Track-D, P2).
Append 1 question to pin down where the timing slipped: the BT trigger or the audio cue.
Put #5 on hold, and state a question asking whether it is 'actual render vs. spec.'

The re-run came back with #4 corrected to Track-D / P2 / "check whether the audio cue offset on the BT combat-start node is 0, or whether the BGM clip itself includes 0.5 seconds of leading silence" and #5 to On hold / "reconfirm with live ops: rendering smaller than spec vs. a request to raise the spec itself". At that point the classification was complete.

What the AI did and what the human did split cleanly here. The AI distributed all five requests in one fast pass and produced a table with no empty cells. The human caught the two cases sitting on a blurry track boundary — the BGM that looks like taste but is a synchronization defect, and the font that looks like a defect but may be taste. Filling five cells without gaps and noticing that two of them are filled wrong are different abilities, and this worked transcript hands each to the side that does it better.


16.2.4 Each Track Demands a Different Hand

Once classification ends, each track enters entirely different follow-up work. They start from the same table but arrive at different places.

Requests classified Track-A (align) are handled by the principle "convey the intent, delegate the expression." My reply to the art lead's color request was not a saturation value but an intent: "This fight is boss phase 1, so tension is the core. I'd rather have pressure than brightness. Within that, saturation is art's call." The moment a designer specifies the saturation value directly, art's autonomy shrinks and accountability for the result blurs. Guarding the boundary between intent and expression is the whole of the align track.

Requests classified Track-D (defect) lead to priority rating and a code connection. The duplicate guild reward (P0) went to the code lead on the spot; the BGM sync issue (P2) was filed to the backlog with a cause-hypothesis question attached. On the defect track, the designer's job is not to fix things but to rate the priority and supply precise input. The line between P0 and P2 is "does this affect user assets or progression right now?" Duplicate rewards touch assets directly, so P0; a 0.5-second BGM lag is unpleasant but blocks nothing, so P2.

Requests classified Track-S (schedule) go into the quarterly roadmap. The publisher's cultural guidelines arrived as a one-liner but in practice decompose into multiple items — depiction of religious symbols, color taboos, text direction, character costume. The key is to answer "we'll review it" the moment it arrives and put the whole thing onto the quarterly plan. For anything carrying an external deadline, however small it looks, lead time is everything; start late and it blows up without exception.

Side by side, the three branches compare like this.

Track-A · Align Track-D · Defect Track-S · Schedule Triage question Taste or intent? Designer's job Convey intent only, delegate expression Triage question User-impacting bug? Designer's job Rate P0–P2, loop in code lead Triage question External deadline? Designer's job Break into items, secure lead time

Classify correctly and the same inbox's five lines scatter cleanly into three separate processing lines. Classify wrong and a defect gets dragged into an alignment meeting and eats the clock, or a schedule item starts late and explodes right before the deadline.


16.2.5 Collaborating by Isolating into a TF

Sometimes requests arrive not as one or two items but as a single mass — the weeks before a publisher review, or a phase like a full combat-system overhaul. In those moments, isolate the work itself into a temporary task force (TF) workspace like 95_BattleTF, and when it ends, promote only the decisions back into the canonical docs. The isolation-and-absorption mechanism, and the operating rule of "hand the art team html only (zero md to learn)," were covered in full in the previous chapter, 16.1.

From the 3-track classification view, one line is enough to add. Collaboration that has grown into a mass is usually a Track-S (schedule) item decomposing across multiple disciplines, so when per-track handling can no longer absorb it, you move it into a container one level up: the isolated workspace. If 3-track classification is the entrance, TF isolation is the room that holds the large mass that came through it.


16.2.6 Common Failures and Remedies

Failure pattern Remedy
Treating every request with the same weight Run the 3-track classification on arrival; ask about external dependency first
Misclassifying a schedule item as align Pin the external-deadline question as triage question number 1
Handling a synchronization defect that looks like taste as align For any "timing/value mismatch," suspect a defect first
Declaring taste that looks like a defect a defect Ask back "is it rendering off-spec?" and put it on hold
The designer deciding the expression on the align track Intent only; delegate expression to the discipline
Starting schedule items late Fold into the quarterly roadmap immediately; secure lead time

Half of this table is mistakes at the classification stage; half is mistakes in handling after classification. Even an accurate classification loses its effect if the per-track hand is wrong. (For pitfalls around TF isolation, promotion, and media, see the pitfall table in 16.1.)


Key Takeaways


Beyond Games. The problem of one-line requests being treated with the same weight just because they landed in the same inbox is not a game thing — it is the everyday reality of every product manager and service designer. The classification that splits incoming requests into the three tracks "align (taste/direction), defect (user-impacting bug), schedule (external deadline)" works unchanged across domains. Say a web service PM's messenger receives, all at once, "make the button color a bit brighter" (align), "payment receipts are being sent twice" (defect), and "privacy-law amendment compliance, deadline 3 weeks out" (schedule): check external deadline → user impact → taste in that order, drop each into the first track that matches, put people on the payment bug immediately, and start the legal change by securing lead time.


Try It Yourself

Minimal path with a web chatbot (no terminal) — the core of this chapter is not the workflow script but the idea: split one-line requests into the three tracks of align, defect, and schedule. That idea reproduces as is in a web chatbot (ChatGPT or Claude on the web), with no CLI, hook, or atom infrastructure. The two steps below are the main line. 1. Collect the day's incoming requests, one line each, in no particular format. Scrape them from messengers, email, memos — anywhere. 2. Paste the prompt below into the chatbot's input box, then paste your collected request list beneath it. This is doing by hand, once, the first-pass classification that request-triangulate used to do. Classify the requests below into Track-A (align) / Track-D (defect) / Track-S (schedule). Check external deadline → user-impacting bug → taste/intent in that order and lock in the first track that matches; if none match, put it on hold and give me 1 question to ask back. Output: | Track | Rationale, 1 line | Next action, 1 line |. [Paste request list] Then verify just two cells of the output table yourself — if a "timing/value mismatch" got classified as align, suspect a synchronization defect; and if a felt complaint like "X is too small/slow" got declared a defect, ask back "is it rendering off-spec?" and demote it to hold. Bring in scripts and workflows only when this classification is second nature and the daily batch becomes too heavy to do by hand.

setup. Gather incoming external requests in one place (a channel or a document). Write down the three track definitions, one line each — align (taste/intent), defect (user-impacting bug), schedule (external deadline).

prompt. Throw the collected batch at the AI and pin the triage order.

Classify the requests below into Track-A (align) / Track-D (defect) / Track-S (schedule).
Check external deadline → user-impacting bug → taste/intent in that order and lock in the first track that matches;
if none match, put it on hold and give me 1 question to ask back. Output: | Track | Rationale, 1 line | Next action, 1 line |.
[Paste request list]

verify. Verify two spots in the output table yourself. (1) If a "timing/value mismatch" was classified as align, suspect a synchronization defect. (2) If a felt complaint like "X is too small/slow" was declared a defect, ask back "is it rendering off-spec?" and demote it to hold. Catch just those two boundary cases by hand and you can trust the rest.

Solo Scale-Down

If you are a solo developer with no team and no TF, keep the tracks and swap only the inputs. Collect store reviews, Discord reports, and beta tester notes in one document, and batch-classify them once a week with the prompt above. For align (taste): accept it if it doesn't conflict with your vision. For defects (bugs): handle them that week. For schedule (store review, event deadlines): enter them in your calendar with the lead time attached. For running an isolation folder during a focused cleanup period, follow the Solo Scale-Down in 16.1.

16.3 One Decision, Three Packages — Framing Deliverables by Discipline

The 95_BattleTF meeting room. On the afternoon we locked the guild attendance reward at resources +5, I sent that same single decision down three channels: a spec in markdown to the design team's channel, a single data column to the programming team, a one-screen html to the art team. Replies came back from all three almost at once. The programming lead asked where the trigger fired, the art director asked whether the attendance button's position matched the 06_UI guide, and the animator said nothing. It was the same decision, and what the three of them saw was entirely different.

This chapter is the record of turning that "seeing differently" from an accident into a design. Packaging one decision differently for each discipline — that is framing.


16.3.1 Five People Read the Same Decision Differently

Five audiences attach themselves to a single guild attendance reward decision. Even when they read the same sentence, each picks out only their own territory and lets the rest slide. The accidents happen where things slide.

Audience What They Read Closely What They Instinctively Skip
Code lead Data columns, interfaces, trigger timing Color, narrative, presentation
Art director Screen layout, components, style guide Data integrity, triggers
Sound director Action triggers, mood, duration Data details
Animator Motion, timing, state transitions Visual tone, numbers
QA Acceptance criteria, risks, edge scenarios Implementation internals

The problem is not the amount of information but the way it is exposed. Put one thick spec on five desks identically, and the five will each open a different page and close a different page. Framing takes that opening out of chance and arranges it deliberately.

Below is a framing matrix showing how one decision changes clothes as it crosses discipline boundaries.

Decision: attendance reward = resources +5 95_BattleTF / single fact Design → markdown Full intent, rules, rationale Includes context to study spec_guild_attendance.md Art → html One screen: layout, components Zero md to study (handoff only) guild_screen_v3.html Programming → data Column, interface, trigger Verification lint items specified guild_table 1 row Invariant facts (what all three packages must preserve) Value = +5 · Timing = first login of the day · Scope = every guild member Packages may differ — if these three values diverge, framing has failed

The packaging differs per audience, but the invariant facts laid underneath — value, timing, scope — must not waver in any of the packages. The craft of framing is not "showing things differently"; it is "showing things differently while preserving the same thing."


16.3.2 Worked Transcript — One Decision into Three Packages

Hand-molding the framing from scratch every time is enough of a burden that it eventually stops happening. So I throw in one decision and receive draft packages for every discipline in a single pass. Below is the actual exchange over the guild attendance reward decision, transcribed as is, unpolished.

Full Prompt

[Decision facts — confirmed in 95_BattleTF]
- Guild attendance reward = guild resources +5
- Trigger: once, when a guild member logs in for the first time that day
- Scope: every member of that guild, individually
- UI: resource total displayed at the top of the guild screen + an "Attendance" button

Package this one decision for three disciplines.
For design, markdown covering intent, rules, and data impact; for art, the screen spec going into the html,
centered on layout, components, and style guide references (no mention of data or triggers);
for programming, one data row, the interface, the trigger, and verification items.
Bake the invariants (+5, once on first login, every guild member) identically into all three packages,
and don't make up unknown values — mark them [needs confirmation].

Claude's Raw Output (No Summarizing, No Cleanup)

1) Design → markdown

## Guild Attendance Reward
- Intent: convert one login per day into a guild-level cooperative resource to lift retention.
- Rule: guild resources +5 when a guild member logs in for the first time that day. Once per day. Accumulates at the guild level.
- Data impact: guild_resource_total increases. Resource inflation needs monitoring [needs confirmation: cap undecided].

2) Art → screen spec (content for the html)
- Resource total displayed in the top area of the guild screen (progress_bar + number badge).
- "Attendance" button on the right. State changes on press follow the button state rules in the 06_UI guide.
- Style: conforms to the 06_UI guide. Colors: [needs confirmation: guild theme color tokens unspecified].
- (Data values and triggers intentionally excluded from this package)

3) Programming → data/interface
- Data: add a column to guild_table — guild_resource_total (int, default 0)
- Interface: GuildResourceManager.add(guild_id, amount)
- Trigger: call add(guild_id, 5) once from the daily first-login event hook
- Verification: integrity_check_guild_resource (no negative values / block calls beyond once per day)
- Concurrency: duplicate grants when the same guild member logs in on multiple devices at once [needs confirmation]

Human Verification and Rejection

I did not use the raw output as is. I caught three things.

Re-request

Revise the art package:
- Remove every external-document reference phrase like "follows the 06_UI guide."
- Describe the visual differences of the button's three states — pressed/idle/done — directly in the screen spec.
- Make it self-contained, on the assumption that the art team works from this one page alone, with no jumps to other documents.

Drop the 2 [needs confirmation] items from the programming package's deliverable,
and pull them out into an "Items for 95_BattleTF re-confirmation" block at the top.

That single rejection and re-request put the deliverables into a form each of the three disciplines could pick up and use right where they sat. The AI molded three draft packages and even flagged the gaps, but the scissor work — deciding what to trim from which package, pulling the external references out of the art package and the unconfirmed items out of the programming package — stayed in my hands to the end. The core judgment in framing lies on the side of exclusion, not inclusion.


16.3.3 Three Framing Methods and When They Pay Off

The method splits on where you put the packaging. Which of the three to use is decided by the size of the spec and your operational stamina.

(1) Per-audience summaries inside one document. Append per-discipline summary sections after the body. Five people share one file, but each reads only their own section.

## Per-Audience Summary

### Code (Implementation)
- Data: guild_table.guild_resource_total (int)
- Interface: GuildResourceManager.add(guild_id, amount)
- Trigger: once on first login of the day
- Verification: integrity_check_guild_resource

### Art (Visual)
- Screen: resource total at the top of the guild screen + attendance button
- Components: progress_bar, badge, button (3 states)
- Priority: this milestone

### QA (Verification)
- Acceptance criteria: guild resources +5 applied after attendance; block anything beyond once per day
- Risks: resource inflation, duplicate grants on multiple devices

(2) Separate deliverables per audience. One body document, with per-discipline files split off on their own. The 95_BattleTF practice of sending the art team only the html and never the md is this method in production form — even for the same decision, the medium itself differs by discipline.

spec_guild_attendance.md     — design body (full context)
guild_screen_v3.html         — art (html only, zero md to study)
guild_table 1 row + add()    — programming (data/interface)
qa_guild_attendance.md       — QA (acceptance criteria, risks)

It suits large specs, and each medium goes straight into the discipline's own tools. In exchange, when one decision changes, several deliverables must be fixed together, so the operational burden is heavy.

(3) A wikilink graph. Put only the per-discipline entry points into the body as links, and let everyone explore down their own branch.

[[spec_guild_attendance]]
   ├── [[code_guild_table]]
   ├── [[ui_guild_screen_v3]]
   └── [[qa_guild_attendance]]

The costs and payoffs of the three methods are as follows. Among the figures below, the "effect" values are the author's estimate (unverified); trust only the direction and the relative ratios.

Method Cost When It Pays Off
(1) Per-audience summaries Around +30% body length Pays off immediately on almost every spec
(2) Separate deliverables Operating N sets of deliverables Pays off only when the spec is large and the media differ by discipline
(3) Wikilink graph Upfront investment in graph infrastructure Pays off once specs accumulate and the graph itself is an asset

For most specs, (1) is the right fit. It costs the least and pays off the fastest. Use (2) only where the media have already split, as with the art html, and switch on (3) once enough specs have piled up for the link graph to earn its exploration value.


16.3.4 Fix the Audience at Five

Redefining the audience for every spec means re-molding the framing every time. So I fix the codes.

Audience Code Domain
code Code, systems, data
art Art, visuals, UI
sound Sound, audio
anim Animation, motion
qa QA, verification

These five are the internal operating standard. External audiences such as outsourcing partners and legal are handled separately, outside this standard. Fixing the count at five means the audience definitions don't have to be rewritten each time framing is handed to an LLM, and a checklist can catch any audience that got dropped.


16.3.5 Automation and Its Pitfalls

Hand-writing five discipline summaries for every spec eventually means not writing them at all. So I bundled the flow like this.

flowchart LR
    A[Designer: write the decision facts] --> B[LLM: draft packages for 5 audiences]
    B --> C[Designer: judge reject/reinforce/keep]
    C --> D{Gap found?}
    D -- Yes --> E[Send back to 95_BattleTF for re-confirmation]
    D -- No --> F[Final spec + per-discipline framing]
    E --> A
    classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764;
    classDef human fill:#fde68a,stroke:#b45309,color:#000;
    classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d;
    class A,C,D human;
    class B ai;
    class F pass;

The designer writes only the decision facts, the LLM produces the five draft packages, and the designer judges each piece: reject, reinforce, or keep. When a gap ([needs confirmation]) surfaces, it is not handled in framing but sent back to the decision stage — because framing is a tool for carrying a settled decision, not a tool for plugging the holes in one.

Here are the four pitfalls this cycle steps into repeatedly, each with its remedy.

Pitfall Symptom Remedy
Duplicated information The same content repeats across body and summaries, raising the upkeep burden State it once in the body; summaries carry only what differs
Missing information A value critical to one discipline drops out entirely Check for omissions against the fixed five-audience checklist
Body ignored People read only the summary and let the body's context slide End each summary with "rationale is in the body"
Medium confusion Sending md to art, imposing a study burden Fix the discipline-medium principle (art = html)

Automation lowers the writing burden to around five minutes per spec, but the reject-reinforce-keep judgment does not get automated along with it. That judgment is the human's seat.


16.3.6 Measurement — With Framing On

The following compares before and after introducing framing on Project A, which I run. The absolute figures are the author's estimate (unverified); what to trust is the direction of change and the relative ratios.

Item Without Framing With Framing Direction
Per-discipline misreading incidents 15–20 per quarter 3–5 per quarter Sharp decrease
Time an audience spends reading a spec 15–30 minutes 5–10 minutes (own section only) Decrease
Decision → work start 1–2 days 4–8 hours Shortened
Cross-discipline interpretation conflicts 8–12 per quarter 2–3 per quarter Decrease
Spec writing time 1–2 hours 1.5–2.5 hours (LLM-assisted) Slight increase

Writing the spec itself gets a little longer, because the per-discipline packaging is layered on top. But the discipline work cycles that follow get shorter, so the total time from decision to work start goes down. That trade-off is the core case for adopting framing. For a team that finds LLM-assisted review burdensome, the safe order is to first establish handwritten five-audience summaries under method (1), then layer automation on top.


Beyond Games. Framing — packaging one decision differently per audience while preserving the invariant facts (value, timing, scope) everywhere — carries straight over to announcement and release communication in any organization, not just games. Say you decide one thing: raise the subscription fee to 9,900 won (about $7), effective July 1. For the development team it is packaged as data, the billing table column and the effective date; for the design team, a single announcement banner screen; for customer support, a response script for anticipated inquiries. The three packages all differ, but the moment the three numbers — 9,900 won, July 1, all new and existing subscribers — diverge in any package, a customer dispute erupts on the spot.


16.3.7 Try It Yourself

setup

prompt

[Decision facts]
(one line each for value, timing, and scope)

Package this decision for the relevant disciplines among code, art, sound, anim, and qa.
In each package, drop the information that discipline doesn't care about, but bake the invariants (value, timing, scope) identically into every package.
Make the art package self-contained on its single page with no references to other documents, and don't make up unknown values — mark them [needs confirmation].

verify

Solo Scale-Down

If you work alone, cut the audience down to two — "future me" (implementation) and "the reviewer" (QA). Write one line of decision facts, ask the LLM to "split this into an implementation memo and a review checklist," and then just check that the key numbers match across the two. Even with only two audiences, the skeleton of framing — package the same decision differently while preserving the invariants — works exactly the same.


Key Takeaways

Part 17 · Meeting Notes

17.1 Why Meeting Notes Are the Biggest Pain

Five minutes after the meeting ended, the whiteboard in the meeting room still had the writing on it: "Combat hit detection goes client-side first. But server validation priority moves to the next sprint." A conclusion five people spent 30 minutes reaching. Everyone nodded, and someone took a photo.

Three weeks later, the same five people gathered in the same meeting room. The first line of the agenda read: "Combat hit detection — client-side first vs. server validation. Decision needed." Nobody remembered the conclusion from three weeks earlier. The whiteboard photo was somewhere in someone's camera roll, and that someone was on vacation that day. Another 30 minutes were spent. This time the conclusion came out the opposite way.

That is the entire reason meeting notes are the biggest pain. Decisions do get made in meetings. But those decisions fail to propagate to the next meeting, the next document, the next build. The decision happened; the propagation did not. This chapter is the story of reconnecting that broken link with data.


17.1.1 Meeting → Decision → Execution: Where It Breaks

The personal R&D system I run is split across 17 documents — an atom naming standard, relationship-map automation, a Layer mapping guide, JIT injection infrastructure, and so on. The single document that has absorbed the most time among them is the meeting-notes improvement plan. Its weight is comparable to the other 16 combined.

At first this puzzled me. Meeting notes — don't you just write down what was said? But when I measured, the pain was located not in writing meeting notes but in what happens after them. Decisions were clearly made in the meeting. The problem was that the moment we walked out the meeting room door, it all evaporated — whose responsibility the decision was, on what grounds, by when, and what it led to.

Drawn as a single scene, the break looks like this.

flowchart LR
    M[Meeting
30-minute discussion] --> D[Decision made
everyone agrees] D -.broken.-> X[Re-tabled at
the next meeting] D -.broken.-> Y[Never reflected in documents] D -.broken.-> Z[Missing from the build] X --> M style D fill:#ffe6cc,stroke:#d79b00 style X fill:#f8cecc,stroke:#b85450 style Y fill:#f8cecc,stroke:#b85450 style Z fill:#f8cecc,stroke:#b85450

The dotted lines are the broken propagation. The decision (orange) was made, but it flows to none of the three branches (red). A decision that fails to flow comes back as a meeting three weeks later. That loop — the arrow bending upward and returning to the meeting — is the body of the pain.

When propagation breaks, four things happen at once.

You lose the history of decisions. "Why did we decide that?" gets answered with "Nobody remembers, so let's schedule another meeting." Repeat meetings multiply. The same agenda item gets re-tabled every quarter. New joiners cannot pick up context. Because the accumulation of decisions is invisible, you have to explain it 1:1 every single time. And AI assistance goes limp. With no context, the answers stay generic. If your meeting notes are scattered, you cannot hand the AI "how our team decided this item before," and the AI returns the internet's average.

All four grow from the same root. Decisions are treated as memos, not as data. Memos evaporate; data flows.


17.1.2 Treat Meeting Notes as a Decision Database, Not a Deliverable

Here you have to flip your perspective once. If you see meeting notes as "the deliverable of a meeting," then writing them down and filing them completes the mission. A filed meeting note is like a memo slip on your desk: visible that day, gone who-knows-where the next week.

If you see meeting notes as a "decision database," it becomes a completely different job. The asset is not the meeting note itself but the decisions extracted from it, and those decisions must be searchable, referenceable, and propagatable. The meeting note is merely the ore vein you draw decisions up from.

Enforcing this shift in code is the backbone of all of Part 17. My system carries one atom that props up this shift. Its name is decision_summary_not_clickup_mirror. Spelled out, it is the principle that "the decision summary in meeting notes is not a mirror of ClickUp (our task tracker)."

Why this atom was needed strikes the dead center of the pain. Many teams simply copy the decision slot of their meeting notes onto the task board. Then the "to-do" survives, but the "why we decided that (the rationale)" disappears. A task tracker holds what to do, not why it was decided. This is exactly why the meeting repeats three weeks later. The tasks were closed, but with no rationale, when someone asks "wait, why did we decide to do it this way?" there is no one who can answer. So the decision summary must not become a mirror of the tracker; it has to be an independent asset that carries its rationale. The atom name itself is that line you do not cross.


17.1.3 Split a Decision into Four Fields

The difference between a decision that propagates and a decision that evaporates lies in structure. A decision that evaporates is a single sentence: "we go client-side first." A decision that propagates decomposes into four fields.

1 decision = 4 fields decision What was decided "Hit detection client-side first" owner Who is accountable Team member A ([MISSING] if empty) rationale Why it was decided "Perceived responsiveness first; accept cheating risk" follow_up What it leads to "Server validation task next sprint" If owner is empty, the pipeline flags [MISSING] — enforcing the line of accountability for propagation

Of the four fields, the most important is owner. A decision without an owner is nobody's job, and a decision that is nobody's job does not propagate into execution. That is why my extraction pipeline does not just move on when owner is empty — it explicitly flags [MISSING]. The very fact that the line of accountability is empty gets pulled up to the surface.

rationale is where the decision_summary_not_clickup_mirror principle lives. Without the rationale, the meeting repeats three weeks later. follow_up is the bridge that carries the decision into actual execution. If this field is empty, the decision remains only a decision and never reaches the build.


17.1.4 The Extraction Pipeline — Drawing Decisions Up from Meeting Notes

A person could fill these four fields by hand every time, but then the enforcement is weak. My system uses a pipeline that automatically extracts decisions from meeting notes and flags missing fields. Three scripts connect in series.

flowchart TD
    R[Meeting note .md
standard template] --> L[meeting_lint.py
template validation] L -->|pass| P[decision_parser.py
extract 4 fields] L -->|reject| R P -->|no owner| MISS["[MISSING] flagged"] P -->|all 4 fields present| PEND[pending atom
1-week verification hold] PEND --> PROM[promote.py
weekly promotion] PROM --> ATOM[official atom
+ JIT manifest registration] style L fill:#dae8fc,stroke:#6c8ebf style P fill:#d5e8d4,stroke:#82b366 style PROM fill:#e1d5e7,stroke:#9673a6 style MISS fill:#f8cecc,stroke:#b85450

The first stage, meeting_lint.py, checks whether the meeting note follows the standard template: is the frontmatter there, are the four slots (agenda/decisions/actions/next meeting) filled in? A note with a broken template is rejected here and returned to its author. An automatic parser can only process input whose format is enforced, so this lint acts as the entry gate for the whole pipeline.

The second stage, decision_parser.py, is the core. It reads the decision slot and decomposes it into the four fields (decision/owner/rationale/follow_up). When it cannot find an owner, it does not discard the decision — it flags it as [MISSING]. Quietly letting an ownerless decision pass is the most dangerous thing you can do.

The third stage is that an extracted decision does not immediately become an official asset; it waits one week in pending status. This verification period is a reversible gate. If within that week it turns out that "this was a discussion, not a decision" or "the rationale was wrong," it gets discarded. Then promote.py moves only the decisions that survived the weekly review into the official atom folder and registers them in the JIT manifest. From the next session on, a registered decision is automatically injected into related work. At last, the decision begins to flow.

The boundary between reversible and irreversible sits right here. Up to a pending discard, everything is reversible. But once promote completes and the decision has propagated into other documents, data sheets, and builds, from that point it is irreversible — teammates' understanding changes and dependent decisions stack on top of it. So all review must finish just before promote, that is, inside the pending reversible window.


17.1.5 Worked Transcript — Running One Broken Meeting Note Through

Speaking only in the abstract does not land, so here is the record, copied as is, of actually feeding one broken meeting note into the pipeline. The input is a meeting note whose format is about half collapsed.

Input — 2026-06-02-battle.md (bad format)

---
type: meeting_note
date: 2026-06-02
category: battle
---

## Agenda
1. Hit detection location (client vs. server)

## Decisions
- Hit detection will be applied client-side first. The prevailing view:
  perceived responsiveness matters more than the server round trip. Accept the cheating risk.
- Server validation goes to the next sprint as a separate task.

## Action Items
- [ ] Design server validation

This note did record the decisions, but owner appears nowhere. Who takes server validation is blank. The next-meeting schedule slot is missing entirely. I fed it to the parser and watched what came out.

Prompt (the full LLM-assist prompt attached to decision_parser.py)

Read the "## Decisions" slot of the meeting note below and extract
4 fields for each decision: decision, owner, rationale, follow_up.
- If the owner cannot be identified in the text, you must write "[MISSING]".
  Do not fill it in by guessing.
- For rationale, quote only the grounds written in the text. If none, "[MISSING]".
- Link follow_up to the action items or any next-sprint mentions.
Output only a JSON array.

[meeting note body pasted here]

Claude's raw output (as is, not summarized)

[
  {
    "decision": "Apply hit detection client-side first",
    "owner": "[MISSING]",
    "rationale": "Perceived responsiveness matters more than the server round trip. Cheating risk accepted",
    "follow_up": "Proceed with server validation as a separate task in the next sprint"
  },
  {
    "decision": "Split server validation into a separate task for the next sprint",
    "owner": "[MISSING]",
    "rationale": "[MISSING]",
    "follow_up": "Linked to the action item 'Design server validation' (owner unassigned)"
  }
]

I stopped here and reviewed the output. Both decisions came back with owner as [MISSING]. Correct — there was no owner anywhere in the note, so not filling it in by guesswork is the right behavior. The second decision's rationale is also [MISSING], and in fact the text contained no grounds for the second decision. That means the parser did not make anything up, which earns trust.

One spot was ambiguous, though. The first decision's follow_up — "server validation in the next sprint" — overlaps with the body of the second decision. The parser saw the boundary between the two decisions a little blurrily. This is where a human steps in.

Human veto + re-request

The first decision's follow_up duplicates the body of the second decision.
Narrow the first decision's follow_up to "create the server validation task,"
and move the specifics of server validation over to the second decision.
Also, since both decisions have owner [MISSING],
prepend a warning line at the very top of the output saying that owner
confirmation is required before this meeting note goes up to pending.

Claude's re-output (raw)

⚠️ Warning: both decisions have owner [MISSING]. Owner confirmation
   from the meeting facilitator is required before pending promotion.

[
  {
    "decision": "Apply hit detection client-side first",
    "owner": "[MISSING]",
    "rationale": "Perceived responsiveness first; cheating risk accepted",
    "follow_up": "Create server validation task (details in decision #2)"
  },
  {
    "decision": "Split server validation into a separate task for the next sprint",
    "owner": "[MISSING]",
    "rationale": "[MISSING]",
    "follow_up": "Link to the 'Design server validation' action item"
  }
]

What this transcript shows is the heart of Part 17. The parser drew the decisions up, but it did not hide the fact that the line of accountability — owner — was empty. [MISSING] got stamped twice, and that went back to the meeting facilitator as a signal: "confirm the owners." For a decision to propagate into execution there must be an owner, and when there is none, the system pushes that fact up to the surface. For one broken meeting note to pass this gate, a human has no choice but to fill in the owner. The first knot of propagation gets tied right here.

For reference, the ⚠️ in the output above is just a console warning line, not part of the document format. The meeting note itself still remains clean four-slot Markdown.


17.1.6 Pruning the Side Branches and Standing Up the Spine

Part 17 was originally designed as six chapters (motivation, extraction, categories, captions, sync, AI assistance) and was consolidated into four. Image captions and sync were side branches of the meeting-notes problem, and it was right to bundle the biggest pain — decision propagation — into one chapter and put it up front. The biggest pain (meeting → decision → execution propagation) was pulled up to §17.1, and the pipeline that resolves it (meeting_lint → decision_parser → promote) was placed right after.

Seeing meeting notes as a database, splitting decisions into 4 fields, the gate that flags [MISSING] when owner is empty, and the one-week reversible verification in pending — these four reconnect the broken propagation. The pain of "the decision was made but it does not flow" dissolves once you shape the decision into a form that can flow.


Key Takeaways


Beyond Games. The pain of "a decision was made in the meeting, but it fails to flow into the next meeting" repeats every week in the meeting rooms of every workplace, not just game development. Writing a decision down not as a one-sentence memo but split into four fields (what, who, why, next action), and surfacing an empty owner as [MISSING], transfers to any meeting as is. For example, if your weekly marketing meeting agreed that "the next campaign centers on Instagram," attach owner (who executes it), rationale (why Instagram — last quarter's conversion-rate evidence), and follow_up (draft the budget). Three weeks later, the question "who was supposed to do that?" never comes up again.


Try It Yourself

Minimal path with a web chatbot (no terminal) — The core of this chapter is not the scripts but the idea: split decisions into 4 fields so they can flow. That idea reproduces as is with only a web chatbot (ChatGPT or Claude on the web), without CLI, hooks, or atom infrastructure. The three steps below are the main path. 1. When a meeting ends, copy the meeting notes (or your meeting memo) as is. No template needed. 2. Paste the prompt below into the chatbot's input box, then paste the meeting notes underneath it (where [meeting note body] goes). This does by hand, once, what decision_parser.py was doing. From the meeting note below, extract 4 fields per decision as a table: decision (what), owner (who is accountable), rationale (why), follow_up (next action). - If the owner cannot be identified, you must write "[MISSING]". No guessing. - For rationale, only grounds written in the text. If none, "[MISSING]". [meeting note body] 3. For every cell stamped [MISSING] in the output table, check with the meeting facilitator and fill it in. Append the completed tables, in date order, to a single document like decisions.md — that document is itself your decision database. For search, in-document find (Ctrl+F) is plenty. Bring in scripts, atoms, and JIT only when this habit has piled up enough that search becomes a burden.

setup (infrastructure version — once the minimal path above feels familiar) - Settle on one standard meeting-note template: frontmatter (type/date/category) + 4 slots (agenda/decisions/actions/next meeting). - Set up three scripts: meeting_lint.py for template validation, decision_parser.py for decision extraction, and promote.py for promotion (at first, lint and parser alone are enough). - Make the 4 decision fields explicit: decision, owner, rationale, follow_up. Put a rule into the parser that forces [MISSING] when owner is empty.

prompt (the LLM-assist prompt attached to decision_parser)

Read the "## Decisions" slot of the meeting note below and extract 4 fields
for each decision: decision, owner, rationale, follow_up.
- If the owner cannot be identified, you must write "[MISSING]". No guessing.
- For rationale, quote only grounds written in the text. If none, "[MISSING]".
- Link follow_up to action items and next-sprint mentions.
Output only a JSON array.
[meeting note body]

verify - Check that every decision in the output has an owner filled in. If there is even one [MISSING], request owner confirmation from the meeting facilitator before anything goes up to pending. - Check that rationale does not invent grounds absent from the text (if there are none, [MISSING] is the correct result). - Leave decisions in pending for one week, then promote only the ones that survive the weekly review to official atoms with promote.py.

17.1.7 Solo Scale-Down

If three scripts feel like too much, shrink it down like this. Unify only the meeting-note template into 4 slots, and when a meeting ends, tear off just the decision slot and run it through an LLM once with the prompt above. Only for decisions whose owner comes back [MISSING], write in the owner on the spot. Even this alone, with no automation, blocks the single most common propagation break: a decision with no owner. Add lint and promote when the notes pile up and you actually need search.

17.2 An Extraction Pipeline That Mines Decisions from Meeting Notes

Wednesday morning, the moment I got to work, a notification popped up in the team messenger. "We did decide last week to expand inventory to 30 slots, right? And who was supposed to fix the data sheet?" Nobody in the thread can answer. The meeting notes definitely exist — in a folder somewhere. Open them and the agenda items and discussion are packed in tight, but "so what did we decide, and who owns it" is dissolved somewhere between the sentences. So at the next meeting, the same agenda item gets brought up again from scratch.

This chapter is the story of the machine that fills that three-day gap. One meeting note goes in; it passes a format check, the four decision fields get extracted, ownerless decisions get tagged [MISSING], a candidate file is created, and a week later, after review, it becomes an asset that is injected automatically. Human hands touch only the two ends: the entrance, where the meeting note is written, and the exit, the once-a-week review.


17.2.1 The Full Pipeline Flow

First, the whole thing in one picture. Each box is either a small script or a human judgment call. Only two boxes are handled by hand; the rest flows automatically.

flowchart TD
    A["Meeting held"] --> B["Write meeting note in standard format
(human)"] B --> C{"meeting_lint.py
format check"} C -->|violation| B C -->|pass| D["decision_parser.py
extract 4 decision fields"] D --> E{"Has an owner?"} E -->|no| F["[MISSING] flagged
request owner assignment"] F --> B E -->|yes| G["Create pending atom
candidate file"] G --> H{"Weekly review
(human)"} H -->|promote| I["promote.py
official atom + JIT registration"] H -->|discard| J["Preserve discard history"] H -->|hold| G I --> K["Auto-injected from the next session"] style B fill:#e8f0ff,stroke:#3366cc style H fill:#e8f0ff,stroke:#3366cc style F fill:#fff0e8,stroke:#cc6633 style J fill:#f0f0f0,stroke:#999999

Only the two blue boxes (writing the meeting note, the weekly review) are human; everything else is a script. The orange box (the [MISSING] flag) is where the automated check calls a human back. When a decision has no owner, the pipeline doesn't simply stall — it sends the decision back to the note-writing stage until someone is made responsible. That is the core design of this pipeline: don't pass over blanks quietly, flag them loudly.

The overall asset folder structure is laid out like this.

meeting_pipeline/ scripts/ meeting_lint.py format and required-section check decision_parser.py extract 4 decision fields + flag owner [MISSING] promote.py pending → official atom + JIT manifest update meetings/ 2026-05-18_battle_tf.md standard-format meeting note (input) atoms/pending/ meeting_decision_2026-05-18_D1.md candidate (awaiting 1-week verification)

17.2.2 Step 1 — A Lint That Enforces the Format

For extraction to work at all, the meeting note has to be in a shape a machine can read. If there is no "## Decisions" section, or the decisions are blended into a paragraph of running prose, the parser can't pull out anything. So the very first thing in the pipeline is a format check. What meeting_lint.py does is simple: is the required frontmatter there, are the required sections there, and are the decision slots filled in D1, D2 format?

# meeting_lint.py skeleton
REQUIRED_FRONTMATTER = ["type", "date", "category", "attendees"]
REQUIRED_SECTIONS = ["## Agenda", "## Decisions", "## Action Items", "## Next Meeting"]
ALLOWED_CATEGORIES = ["art", "battle", "daily", "issue", "review"]

def lint(meeting_note_path):
    fm, body = parse_markdown(meeting_note_path)
    errors = []
    for key in REQUIRED_FRONTMATTER:
        if key not in fm:
            errors.append(f"Missing frontmatter: {key}")
    if fm.get("category") not in ALLOWED_CATEGORIES:
        errors.append(f"Invalid category value: {fm.get('category')}")
    for section in REQUIRED_SECTIONS:
        if section not in body:
            errors.append(f"Missing section: {section}")
    if "## Decisions" in body:
        block = extract_section(body, "## Decisions")
        if not any(l.strip().startswith("- D") for l in block.split("\n")):
            errors.append("Decision slots empty (D1, D2... format required)")
    return errors

I wire this check into a pre-commit hook for meeting notes. Break the format and the commit itself is blocked. Leave it as a recommendation and people quietly skip it on a busy day, and a format skipped once collapses the following week. Stay blocked for a week or two and the format becomes second nature. But if it's too strict, people start postponing the meeting notes themselves — so the realistic way to run it is to clean up the false positives once, after the adjustment period.


17.2.3 Step 2 — A Parser That Mines the Four Decision Fields

From a meeting note that passed the format check, decision_parser.py reads the decision slots. There are exactly four things to extract from each decision: what was decided (decision), who is responsible (owner), why it was decided that way (rationale), and what to do next (follow_up). These four fields are what turn a decision into an asset. Especially owner. A decision without an owner isn't a decision; it's wishful thinking. So when owner is empty, the parser doesn't quietly leave a blank — it writes in [MISSING] and raises a flag.

From here to the end, I'll follow one meeting note all the way to becoming an asset, without skipping a single line. It's one continuous example, from input to atom promotion.

================ Input: meetings/2026-05-18_battle_tf.md ================
---
type: meeting
date: 2026-05-18
category: battle
attendees: [Minsoo Lee, teammate_a, teammate_b]
related_atoms: [combat_global_cooldown_constant]
---
## Agenda
- Unify the combat global cooldown (GCD) value
- Whether healing skills are exempt from the GCD

## Decisions
- D1: Unify the combat global cooldown at 0.5 seconds. (owner: teammate_a) [rationale: 0.5 seconds was the most stable in input-response feel tests against the refgame]
- D2: Exclude healing skills from the global cooldown. [rationale: risk of breaking the healing cycle]

## Action Items
- @teammate_a: apply 0.5 across the cooldown column in the combat data sheet (~MM-DD)

## Next Meeting
- MM-DD 14:00, review of the one-week healing-cycle test results

================ $ python meeting_lint.py meetings/2026-05-18_battle_tf.md ================
[OK] frontmatter 4/4, sections 4/4, 2 decision slots detected. Commit allowed.

================ $ python decision_parser.py meetings/2026-05-18_battle_tf.md ================
[
  {
    "id": "D1",
    "decision": "Unify the combat global cooldown at 0.5 seconds.",
    "owner": "teammate_a",
    "rationale": "0.5 seconds was the most stable in input-response feel tests against the refgame",
    "follow_up": "apply 0.5 across the cooldown column in the combat data sheet (~MM-DD)",
    "source_meeting": "2026-05-18_battle_tf.md",
    "category": "battle",
    "related_atoms": ["combat_global_cooldown_constant"]
  },
  {
    "id": "D2",
    "decision": "Exclude healing skills from the global cooldown.",
    "owner": "[MISSING]",          # ← owner not specified. Parser flags it
    "rationale": "risk of breaking the healing cycle",
    "follow_up": null,             # ← no follow-up action either
    "source_meeting": "2026-05-18_battle_tf.md",
    "category": "battle",
    "related_atoms": ["combat_global_cooldown_constant"]
  }
]
[WARN] D2: owner=[MISSING] — decision with no owner. Pending creation withheld; returned to the meeting-note author.

================ pending created: only D1 passes ================
$ cat atoms/pending/meeting_decision_2026-05-18_D1.md
---
name: meeting_decision_2026-05-18_D1
description: Decision to unify the combat global cooldown at 0.5 seconds
status: pending
type: decision
source_meeting: 2026-05-18_battle_tf.md
owner: teammate_a
category: battle
related_atoms: [combat_global_cooldown_constant]
created: 2026-05-18
---
## Decision
Unify the combat global cooldown at 0.5 seconds.
## Rationale
0.5 seconds was the most stable in input-response feel tests against the refgame.
## Follow-Up Actions
- [ ] @teammate_a: apply 0.5 across the cooldown column (~MM-DD)

================ One week later: weekly review ================
$ python promote.py atoms/pending/meeting_decision_2026-05-18_D1.md
[PROMOTE] → atoms/combat_global_cooldown_constant_decisions/meeting_decision_2026-05-18_D1.md
[JIT] manifest registered: trigger=(전투|쿨다운|GCD|cooldown), atoms 18 → 19
[OK] From the next session, typing "global cooldown" auto-injects this decision.

This one box is the entire pipeline. The spot to watch is D2. The decision content is fine and the rationale is there, but owner is empty. The parser does not let it through. It writes in [MISSING], withholds the pending file, and returns the decision to the author. A few days later, D2 picks up an owner at the "review of the one-week healing-cycle test results" meeting and comes back in. That single bounce at the blank is what keeps "wait, who was supposed to do that?" from ever appearing in the team messenger three days later.

The rule itself — flag any decision without an owner — is pinned down in an atom (decision_summary_not_clickup_mirror, §17.1.2). The task tool may well show a to-do that says "fix the data sheet," but why that to-do exists, and what decision it follows from, survives only in the meeting-note atom.


17.2.4 Step 3 — Let Decisions Sit in Pending for a Week

A decision the parser passes does not become an official atom right away; it waits a week in pending/. Decisions made confidently in a meeting are routinely overturned after a week of actually running them. D2 in the example above sat exactly in that danger zone: the decision "healing is excluded from the GCD (global cooldown)" could flip again if the healing cycle breaks down in the one-week test. pending is the slot that forcibly sets aside time for the ink to dry.

And discards are kept as assets too. If a decision like D2 had collapsed in the one-week test, I wouldn't just delete it — I'd create a discard-history atom.

---
name: meeting_decision_2026-05-18_D2_DISCARDED
status: discarded
discarded_reason: healing-cycle DPS curve collapsed in the one-week test
---
## Original Decision
Apply the 0.5-second global cooldown to healing skills as well.
## Reason for Discard
Healing-cycle DPS dropped in the one-week test and broke the overall balance. Reverted to the exclusion decision.
## Lesson
"Healing excluded from the GCD is the standard" → promoted to the combat_healing_skill_cooldown_exception atom.

The discard history becomes the answer to "didn't we try this before?" at the next meeting. It's the cheapest tool there is for not making the same mistake twice. Discard records do pile up into search noise, though, so they need a quarterly trim: deduplicate and keep only the lessons.


17.2.5 Step 4 — The Weekly Review and Promotion

At a fixed time each week, I go through the pending candidates in one batch. Each one ends in one of three outcomes.

Outcome Handling
Promote Move pending → official atom folder, register in the JIT manifest
Discard Decision overturned → remove from pending, preserve a discard-history atom
Hold Not enough information → extend pending by one week

The review takes about 15 minutes per 10 atoms. Once promotion is decided, promote.py handles the file move and the manifest update in one shot.

# promote.py skeleton
def promote(pending_path):
    fm, body = parse_markdown(pending_path)
    target = ATOM_BASE / f"{fm['related_atoms'][0]}_decisions" / f"{fm['name']}.md"
    move(pending_path, target)
    manifest = json.load(open(JIT_MANIFEST))
    manifest['atoms'].append({
        "name": fm['name'],
        "path": str(target),
        "trigger_regex": build_trigger(fm),   # related_atoms + category keywords
        "description": fm['description'],
        "added": today(),
    })
    json.dump(manifest, open(JIT_MANIFEST, "w"), indent=2)
    log_promotion(fm['name'])

When trigger_regex matches user input in a later session, this decision is injected automatically. In the example above, typing "global cooldown" brings in the D1 decision along with its rationale. This is the point where a decision I used to carry over by hand becomes an asset that surfaces on its own, the moment it's needed.


17.2.6 Measurement — What Changed from Copying by Hand

This is my impression from running Project A, comparing the stage where I had only the standard format with the stage where the pipeline was running. The numbers below are not precise measurements — they are directions and rough ratios as felt during operation, and they include the author's estimates (unverified).

Item Format Only (Manual Extraction) Pipeline Running
Meeting note → decision extraction time 20–30 minutes per meeting Under 1 minute
Share of decisions promoted to atoms 5–10% (no time to organize) 60–80% (full review)
"Didn't we already decide this?" re-meetings 5–10 per quarter 0–2 per quarter
Ownerless decisions Not tracked Immediately visible via [MISSING] flags

The biggest change is the promotion rate. When I organized by hand, more than 90% of decisions evaporated for lack of time. With automation, full review became possible, and the decisions worth keeping survive without exception. The direction is clear; the exact ratios will vary with team size and meeting frequency.


17.2.7 Common Failures and Their Fixes

Pattern Fix
Running lint as a recommendation only Enforce it with a commit hook
Writing the discussion into the decision slot One sentence per decision; rationale in a separate field
Letting an empty owner through Flag [MISSING] + withhold pending and bounce it back
Pending review keeps slipping A fixed slot in the weekly retrospective — even 5 minutes, every week
Not keeping discard history Preserve discards as separate atoms too

These five lines are nearly all of it. Minimize the spots that depend on human willpower, and hand the format and owner checks to the machine — that is where this system finds its stable point.


Key Takeaways


Beyond Games. This structure — one meeting note rides a conveyor of format check → decision extraction → one-week verification → official registration, and humans touch only the entrance (writing) and the exit (the weekly review) — ports to the document operations of any knowledge-work team, not just games. For example, when a consulting team handles client meeting notes: standardize the note format, have an LLM do a first-pass extraction of the "decision / owner / rationale / next action" slots, raise [MISSING] and bounce back anything with an empty owner, and promote only the decisions that have aged a week into the official action tracker. Meeting decisions that evaporated more than 90% of the time when organized by hand become, on the conveyor, subject to full review — and they survive without exception.


Try It Yourself

setup. Put a standard-format template in your meeting-notes folder and wire meeting_lint.py into a pre-commit hook. Make the four frontmatter fields and the four sections required.

prompt. Feed one meeting note to the parser with an instruction like this.

From the ## Decisions section of this meeting note, extract the four fields decision / owner / rationale / follow_up as JSON for each decision. For any decision without an explicit owner, set owner to [MISSING] and collect it on a separate warning line. Do not fill in guesses.

verify. If the output JSON contains any decision marked [MISSING], do not create a pending file for it — return it to the meeting-note author. Create pending candidate files only for decisions whose owner is filled in, and decide promote, discard, or hold at the weekly review one week later.

Solo Scale-Down

If you work alone, three scripts plus a commit hook is overkill. Standardize only the ## Decisions section of your meeting notes, and write each decision as one line: D1: what / owner: me / rationale: why. Once a week, scrape just the decision lines from that week's notes into a single file (decisions.md), and mark [MISSING] yourself on any line with an empty owner so you fill it in the following week. Adding the scripts later, once your hands start to hurt, is not too late. The core is three habits: one line per decision, an explicit owner, and a weekly collection pass.

17.3 Meeting Categories, Captions, and Sync — The Three Axes That Turn Meeting Notes into Assets

The goal of meeting notes is not to pile them up. The goal is for them to still be searchable six months later, to lead to decisions, and to look exactly the same on two different PCs.


Tuesday afternoon. I remembered that in a meeting a year earlier, we had clearly agreed to lower the saturation of a character's outfit by one step. But I could not find the meeting note. When I opened the folder, 200 files like meeting_0413.md, 회의_수정본_final.md, and IMG_2034.png sat there, sorted only by date. No categories, no captions, no consistent naming. The decision was in there somewhere, but the path to reach it was gone.

For meeting notes to become an asset, three things have to work at the same time. Categories create the first entry point for search, captions keep the image half of the notes searchable, and sync ties processing cost to the changed files only, even past 1,000 notes. If any one of the three is missing, meeting notes become a dead pile that only gets heavier as it grows.

In §17.1 and §17.2 I set up the flow that turns meeting notes into an extraction pipeline — meeting_lint.py checks the format, decision_parser.py pulls the four decision fields (decision / owner / rationale / follow_up), reports [MISSING] when there is no owner, collects candidates as pending atoms, and promotes them with promote.py. This chapter covers the three operating standards that keep that pipeline from decaying over the long run.


17.3.1 Categories — The First Entry Point for Search

Meeting notes grow into the hundreds and then the thousands. Material that cannot be searched is not an asset. Categories are the first fork in that search. It is the same as putting labels on office filing cabinets: a cabinet without labels is one nobody ever opens.

On Project A (an MMORPG in development) that I run, we grouped meetings into five categories. The key is to keep them small and orthogonal.

art Visual/art direction Concept reviews Env tone alignment → captions ↑ battle Combat/balance Cooldowns, DPS Damage curves → atom extraction ↑ daily Routine progress Stand-ups Today's tasks → almost no decisions issue Urgent issue response Build failures Pre-launch incidents → post-hoc cleanup required review Milestones/QA MS sign-off Quarterly retros → summary atoms The five buckets never overlap — one meeting goes in exactly one bucket Grow to six, and "is this art or battle?" stalls a meeting every week

Five is not the right answer for every team. If your project centers on non-combat systems, you adjust — swap battle for system, and so on. The point is not the number; it is the principle of keeping classification decisions small enough that they never stall the meeting itself.

One Meeting, One Category

Meetings that straddle two buckets happen all the time. If a character concept review ended up settling combat motion too, is it art or battle? The rule is exactly one, by primary deliverable. If the concept is the primary deliverable, classify it as art and record the combat motion as a secondary entry in the sub_topic field.

---
type: meeting_note
category: art
sub_topic: [character, battle_motion]
date: 2026-05-18
attendees: [teammate_a, teammate_b, teammate_c, Minsoo Lee]
related_atoms: [character_concept_kim, battle_motion_kim]
confidential: internal
---

sub_topic is only a secondary search filter; it is never used for routing decisions. Routing always operates on the single category value alone. If this single-value rule collapses, promote.py from §17.2 can no longer decide which folder an atom goes to, and per-category statistics stop adding up. Orthogonality is not a matter of tidiness — it is a precondition for pipeline integrity.

Each Category Runs Differently — That Is the Real Value of the Split

The real reason for the five buckets is not search labels. Each bucket is operated differently, and only a clean split lets that differentiated operation fall into place naturally.

art notes carry many attached images, so the caption standard in the next section is mandatory; decisions are visual, so the decision slot holds image references like ![](images/decision_a.png). battle decisions are numbers and rules, so it has the highest rate of automatic atom promotion, and since a one-line decision can cascade into bulk data sheet changes, visualizing the blast radius (the relationship maps from Part 11) matters. daily is supposed to have almost no decisions, and it accumulates fast, so it goes into automatic weekly folders (daily/2026-W21/). issue notes are messy by nature, so cleanup within 24 hours is mandatory and recurrence-prevention atoms are extracted into issue_postmortem/. review notes run long, so a separate 5–10 line summary atom is written and gets auto-cited in the next quarterly retrospective.

Adding a new category is something I treat with great caution. It must occur at least 5 times per quarter, be operated in a way clearly different from the existing five, need its own routing folder, and still be holding at 5 or more occurrences a month later — only when all four conditions pass do I even consider it. In my experience the five have held for over a year, and when candidates like tech_review or external came up, they were ultimately absorbed into sub_topic.

The AI Classifier Must Stay a Backup

Categories are entered by the person writing the note — that is the primary path. Only notes with missing categories, such as material received from outside, get AI-assisted classification. A keyword dictionary catches about 90%, and only the remaining uncertain cases go to an LLM or a human.

When delegating to an LLM, a prompt with hard constraints is the stable choice. Here is the exact prompt I use.

The following is a meeting note. Classify it into exactly one of the 5 categories.

Categories:
- art: visual and art direction
- battle: combat systems and balance
- daily: routine progress sharing
- issue: urgent issue response
- review: milestone and QA reviews

Meeting note:
[full text or first 500 characters]

Response format: one category word only. No explanations, rationale, or hedging of any kind.
Any response that is not one of the 5 categories is treated as a system failure.

When I fed in a meeting note (below is the opening of an art meeting), Claude's raw output looked like this.

Input meeting note: Character K_007 (Scholar) concept v3 review. Feedback that the outfit's color saturation is too high. Agreed to lower it by one step. Combat motion tone to be checked together at the next meeting.

Claude output: art

A clean single word. But when I fed a daily meeting note into the same prompt, this also happened.

Input: The build broke overnight; the cause looks like a data sheet merge conflict. Hotfix first, proper fix to follow.

Claude output: issue

On the surface this was said in a daily stand-up, but Claude read the content and classified it as issue. This is exactly why the classifier must not be the primary path. A human makes the operational call — "this is a build incident that surfaced mid-daily, so it should be split off into a separate issue meeting." The AI only looks at text and stamps a label. The label may be right, but it cannot decide whether the meeting should be split. So humans come first, and the LLM only backfills the gaps.

At the quarterly retrospective I tally meeting counts per category to see where the time goes. The distribution below is the author's estimate (unverified): the absolute counts are illustrative, and only the relative order matches my actual operating experience.

Category Share (estimated) Notes
daily about 1/3 daily and routine, almost no decisions
battle about 1/5 combat task force, twice a week
art about 1/7 art reviews + external meetings
issue low build incidents and the like
review lowest milestones, quarterly retrospectives
Other about 1/5 1:1s, external, and other non-categorized

If issue spikes in a given quarter, improving build and CI stability rises to the top of the priority list. Categories are not just for search — they are a mirror of how the organization spends its time.


17.3.2 Captions — The One Line That Keeps Half Your Images Alive

Half the body of an art meeting note is images. And an image without a caption is like a pile of photos stacked on a desk. On the day itself you remember everything; a month later, only the photos with a one-line note on the back survive.

flowchart LR
    A["Right after the meeting
only attendees understand it"] --> B["1 week later
even the writer recalls only part"] B --> C["1 month later
unclear which decision it relates to"] C --> D["6 months later
effectively discarded · unsearchable"] A -.one caption line.-> E["Even 6 months later
reverse-traceable by decision ID"] style D fill:#fcd6d6,stroke:#d94a4a style E fill:#d6fce0,stroke:#4ad97a

If images are half the meeting note and they cannot be searched, then half of the asset is gone. What keeps that half alive is one line of caption.

The Three Caption Elements

Project A's caption standard ends in three lines.

![](images/2026-05-18_art_review/character_kim_concept_v3.png)

**[Figure 1]** Character K_007 (Scholar) concept v3 — outfit color saturation lowered one step
*Decision: D2 (outfit saturation -10%) | Next action: v4 work (~MM-DD)*

Each of the three elements opens a different search path. The number + one-line description gives the body a way to cite it ("see Figure 1"); the decision ID reference (D2) enables the reverse lookup "images linked to this decision"; the next action leaves a trail to the follow-up work. All three lines take under a minute to write. "Attach immediately" does not mean "write during the meeting." Realistically, you only tidy the decisions during the meeting and fill in the captions within ten minutes after it ends.

File Names and Folders Are the First Entry Point

Just as important as captions are file names, because folders and file names themselves are the first entry point for search.

meeting notes folder/
├── 2026-05-18_art_review.md
└── images/
    └── 2026-05-18_art_review/
        ├── character_kim_concept_v3.png
        ├── env_palette_comparison.png
        └── reference_external_game_a.png

The rule is <topic>_<item>_<version or note>.<ext>, and Korean characters, spaces, and special characters are banned (to prevent path-encoding accidents). IMG_2034.png (zero meaning), 김캐릭터 v3.png (Korean plus a space), final_final_v3_real.png (meaningless versioning), and untitled.png (a deletion candidate) are all anti-patterns. Rather than relying on willpower, it is better to enforce these names by adding a check rule to meeting_lint.py — one extra file-name check on top of the format lint we automated in §17.2 is enough.

External Source Attribution and Confidential Ratings

Meetings often cite external games and art as references. Without attribution, that is a straight line to a copyright incident.

![](images/2026-05-18_art_review/reference_external.png)

**[Figure 3]** Reference image — refgame (Developer Y, 2024)
*Reason for citation: comparing saturation treatment of a similar concept. No direct borrowing.*

State all three: the source (game title, developer, year), the reason for citation, and whether anything was directly borrowed. And because images leak more easily than text, the rating goes in the frontmatter.

confidential: internal   # internal / restricted / external_ok
images:
  - file: character_kim_concept_v3.png
    confidential: restricted
    reason: unreleased character design

internal means company-wide sharing, restricted means the relevant task force and owners only, and external_ok means approved for marketing and external sharing. When the meeting notes are built, output is split by rating, and any image not marked external_ok is automatically blurred in the external-share build. This automatic split is what drives external-share masking accidents to effectively zero.

AI Drafts the Captions, Too

Writing 50 captions by hand for 50 images is a burden. Give the AI the body text and the file names and get a batch of drafts.

The following is a meeting note body plus a list of image files.

[meeting note body]
[10 image file names]

Draft a caption for each image.

Format:
- [Figure N] <description> — <key decision or change>
- *Decision: D? | Next action: ?*

For any image whose grounding cannot be found in the body, mark it as "content unknown — writer confirmation needed".

The last line is the key. Given the same meeting note, Claude wrote captions for the images grounded in the body, but for reference_external_game_a.png it answered like this.

Claude output (excerpt): [Figure 3] reference_external_game_a.png — content unknown, writer confirmation needed. The body does not state the reason for citing this external reference image.

The AI reported what it did not know as unknown. The writer takes that and fills in the citation reason. When the body context alone is not enough, send just the 5–10 key images to a Vision model (image token costs are high, so do not run all of them).

# Apply selectively to the 5-10 key images only — token cost per image is high
response = client.messages.create(
    model="claude-opus-4-8",
    messages=[{
        "role": "user",
        "content": [
            {"type": "image", "source": {"type": "base64", "data": img_b64}},
            {"type": "text", "text": "Describe this image in one line of Korean. No guessing — only what is visible."},
        ],
    }],
)

The writer then shapes that one line into the caption format. There is no need to run Vision on every image; the 5–10 key ones raise searchability plenty.

A year of well-captioned meeting notes becomes a visual development document in its own right. You can trace the visual evolution of character_kim v1 → v2 → v3 by decision ID; filtering on the external_ok rating auto-curates external reporting material; and collecting key images plus captions per discipline produces onboarding material for new team members. Expressing the before/after of caption adoption as the author's estimate (unverified), the direction is this — search success on six-month-old meeting notes rises sharply, the "where did I see this image?" re-asks drop sharply, and external-share masking accidents converge to 0. The absolute numbers will differ by team, but steps 1 and 2 alone (file-name standard + caption format) made the direction unmistakable.


17.3.3 Sync — Only the Changes, Not the Whole

Meeting notes themselves are text files, so git is enough. The real targets of sync are the data derived from them — the pending atom candidates from §17.2, the JIT manifest, category statistics, the decision index (decision_index.json), the caption index, the per-confidential-rating build outputs, and the LLM embeddings for vector search. All of these must react to every meeting-note change.

The problem is that once meeting notes pass 1,000, reprocessing everything each time eats half of the operating budget. It is like stopping the entire production line and remanufacturing every part — when only one part changed.

flowchart TB
    subgraph Full["Full Sync · viable under 100 notes"]
        F1["All 1,000 meeting notes"] --> F2["All derived data
regenerated from scratch"] F2 --> F3["Replace every index and embedding"] end subgraph Inc["Incremental · required past 200 notes"] I1["Detect only the N changed
files via git diff"] --> I2["Regenerate derived data
for those N files only"] I2 --> I3["Partial index update
add · modify · delete branches"] end Full -.past 200 meeting notes.-> Inc style Full fill:#fce7d6,stroke:#d98a4a style Inc fill:#d6fce0,stroke:#4ad97a

Full Sync is simple to implement and carries zero risk of state divergence, so in the early days (under 100 notes) it is actually the safer choice. Full is not a bad method. But its cost scales linearly with the note count, and somewhere past 200 notes that becomes the bottleneck. That is when you switch to Incremental.

Detect Changes with git diff

The first step of Incremental is judging precisely which files changed. File mtime is fast but inaccurate — a mere touch registers as a change. File hashes are content-based and accurate, but weak at distinguishing additions from deletions. My recommendation is git diff. Record the commit hash of the last sync, then process only the files changed since. It catches additions, modifications, and deletions accurately, with the smallest extra state-management burden.

# incremental_sync.py skeleton
def get_changed_files(last_sync_commit):
    result = subprocess.run(
        ["git", "diff", "--name-only", last_sync_commit, "HEAD", "--", "meetings/"],
        capture_output=True, text=True
    )
    return result.stdout.strip().split("\n")

def sync():
    last_commit = read_state("last_sync_commit")
    for path in get_changed_files(last_commit):
        if not os.path.exists(path):
            handle_deletion(path)        # delete atoms, index entries, and embeddings together
        elif is_new(path, last_commit):
            handle_creation(path)        # lint → decision extraction → pending atom → index → embedding
        else:
            handle_modification(path)    # invalidate existing derivatives, then reprocess
    write_state("last_sync_commit", get_current_commit())

There is one more branch that splits the cost most sharply: whether the body of the note changed, or only the frontmatter.

def detect_change_scope(file_path, last_commit):
    diff = subprocess.run(
        ["git", "diff", last_commit, "HEAD", "--", file_path],
        capture_output=True, text=True
    ).stdout
    fm_lines, body_lines = split_diff_by_section(diff)
    return {"frontmatter_changed": bool(fm_lines), "body_changed": bool(body_lines)}

scope = detect_change_scope(path, last_commit)
if scope["body_changed"]:
    full_reprocess(path)          # includes embedding regeneration
elif scope["frontmatter_changed"]:
    metadata_only_update(path)    # zero embedding regeneration

If only metadata like category or confidential changed, there is no need to regenerate the LLM embeddings. Embeddings are usually the single largest chunk of sync cost, so this one branch cuts the bill substantially. Embeddings are cached by content_hash — if the body hash is unchanged, the cached embedding is reused as is, and a frontmatter-only edit results in zero embedding calls.

The direction of the cost difference is clear (the following is the author's estimate, not absolute values). In an operation with about 50 changes per week, Incremental's embedding cost came out dozens of times lower than a weekly full re-embed. As the notes accumulate, Full's cost grows in proportion to the pile, while Incremental's cost stays tied only to the weekly change count — nearly flat, regardless of accumulation. This "accumulation-independent" property is the essential value of Incremental.

Two Safety Nets — Periodic Full Re-Sync and Single-PC Sync

Incremental is fast, but it carries the risk of accumulating divergence. If a small bug drops a single atom, that omission will not heal itself in the next Incremental run. So I bolt on guardrails — Incremental daily, a Partial Full over the most recent week as weekly verification, and a full re-sync monthly to check index and embedding consistency. When the monthly check finds a mismatch, the change-detection logic gets reinforced. That once-a-month pass is the last safety net of long-term operation.

On top of this sits the split-PC setup. I handle meeting notes on two machines, the office PC and the home PC. The rule: sync work runs on one PC only.

Flow Handling
Office PC → git push the office PC owns the sync work (regenerating derived data)
Home PC → git pull update last_sync_commit only; no reprocessing
Both changed, then merge recompute the changed files against the merge result

If both sides sync at once, the last_sync_commit state collides, and that collision silently skews the indexes. Pinning one PC as the sync owner is the simplest rule and the surest defense.


17.3.4 Where the Three Axes Meet in One Pipeline

Categories, captions, and sync are not free-floating standards. The three are bound into one flow on top of the extraction pipeline from §17.2.

When a meeting note is written, category decides the routing in promote.py; the caption's decision ID links to the four decision fields that decision_parser.py extracted; and Incremental sync picks out only the changes to refresh all the derived data built that way. The starting point for operating this chapter is decision_summary_not_clickup_mirror (§17.1.2). Categories open the path to find a decision, captions preserve the decision's visual evidence, and sync keeps that decision asset in the same state on both PCs.

That Tuesday-afternoon helplessness — the state of having clearly agreed on something with no path left to reach it — disappears the moment these three axes are running. category: art narrows the folder, the caption's Decision: D2 lands on the exact decision, and sync shows me that meeting note at home, looking exactly the same.


Beyond Games. The principle that material becomes an asset only when it can be searched, referenced, and synced is not a game-meeting-notes story; it is the shared problem of every working professional who handles documents. The three axes — categories (small and orthogonal), captions (a one-line description for every attached image), and sync (only the changes, never the whole) — stay the same when you swap out the domain. Say a sales team accumulates a year of client meeting material: fix the categories at five or fewer, such as "new proposals / contract negotiation / post-sale support," attach a line like "[Figure 1] Company A second quote — unit price cut 5%" to every quote screenshot, and have the cloud sync pick out only the changed files. That is what lets you answer "why did we cut that price back then?" six months later from a single caption line.


17.3.5 Try It Yourself

setup 1. Define 5 or fewer meeting categories (start from art / battle / daily / issue / review and swap 1–2 to fit your team). 2. Add two checks to meeting_lint.py — that category is one of the defined values, and that image file names match the <topic>_<item>_<version> pattern (no Korean characters or spaces). 3. Add a confidential field to the frontmatter and prepare a state file to record last_sync_commit.

prompt (classification backup for notes with missing categories)

The following is a meeting note. Classify it into exactly one of the 5 categories.
[5 lines of category definitions] / [first 500 characters of the meeting note]
Response format: one category word only. No explanations, rationale, or hedging of any kind.
Any response that is not one of the 5 categories is treated as a system failure.

verify 1. Pick any six-month-old meeting note and try to find it using only the category plus the caption's decision ID. 2. Check that the number of changed files caught by git diff --name-only <last_sync_commit> HEAD matches the number of meeting notes you actually edited. 3. Do not accept AI classification uncritically — have a human re-check the uncertain cases and the "decision that surfaced mid-daily" cases.


17.3.6 Solo Scale-Down

If you are a game designer working alone, shrink it down like this.

Even at solo scale, the one invariant is this — leave a path that reaches the decision. Categories, captions, and sync are just the three pillars holding up that path, and you can make them as thin as your scale allows.


Key Takeaways

Next Chapter Preview

17.4 Turning Meeting Notes into a Decision Database — Five AI Automation Points

Three days before the milestone demo, over lunch, one of the designers sets down a tray and asks: "The quest reward gold going up to 1.5x — we decided that in the meeting, right? Can I put it in the data sheet?" The next seat over answers, "Wasn't that just someone saying, why don't we try it?" The 90-minute recording exists, and so do the two pages of notes someone hammered out on a keyboard. But that record captured "what was talked about" — not "what was decided, who is responsible, and why."

When I lined up the company's 17 R&D documents in order of pain, the largest share belonged to the meeting-notes improvement plan. That surprised me. It wasn't combat balance, and it wasn't the content production pipeline. The sorest spot was one single thing: decisions made in meetings were not propagating into execution.

So I redesigned the meeting-notes system as a decision-tracking database, and over six months of running it myself I verified where AI belongs in that flow and where it does not. This chapter is the map of those five points.


17.4.1 The Five Points Where AI Automation Can Plug In

From the audio recording all the way to updating the decision graph, there are exactly five places in the meeting-notes pipeline where an AI assistant can sit. I do not turn on all five at once. Each spot differs in maturity and in the risk of accidents.

flowchart TD
    A[Meeting audio recording] -->|"Position 1
STT"| B[Raw transcript] B -->|"Position 2
Draft generation"| C[Meeting-notes draft] C -->|"meeting_lint.py"| D[Standard-format meeting notes] D -->|"Position 3
Decision slot enrichment"| E[Decision with all 4 fields] E -->|"decision_parser.py
Position 4 routing"| F[pending atom] F -->|"promote.py
Position 5 relationship extraction"| G[Official atom + graph] style C fill:#ffd9d9,stroke:#c0392b style E fill:#d9f0ff,stroke:#2980b9 style F fill:#d9f0ff,stroke:#2980b9

Red (Position 2) is the most tempting and the most dangerous spot; blue (Positions 3 and 4) are the safe ones I adopted first. The middle of the pipeline — meeting_lint.pydecision_parser.pypromote.py — is deterministic script, not AI. AI enters only the "gaps that need judgment" between the bones of this deterministic skeleton.

One line each on the character of the five points:


17.4.2 The Deterministic Skeleton — Create the Gaps for AI to Fill First

Before talking about AI automation, you have to look at the non-AI script skeleton first. The reason meeting notes can become a decision database lies not in the LLM but in three small Python scripts.

Standard-format meeting notes end with a decisions block. Each decision in the block enforces four fields.

## Decisions

D1:
  decision: Unify the combat global cooldown at 0.5 seconds
  owner: teammate_a
  rationale: At 0.3s, skill-chaining tests showed frequent dropped inputs (transcript 14:22)
  follow_up: Reflect GCD 0.5 in the combo design sheet, by 6/13

decision_parser.py reads this block. Its core behavior is simple — if any of the four fields is empty, it prints [MISSING] and reports it. The hardest stop is reserved for a missing owner: a decision without an owner is "a decision nobody is responsible for," that is, a decision that will not be executed.

$ python decision_parser.py 2026-06-06_combat-sync.md

D1: OK   (owner=teammate_a)
D2: [MISSING owner]  "Consider exempting healing skills from the GCD" — no owner, promotion blocked
D3: [MISSING rationale]  rationale field empty, warning

D2, flagged [MISSING owner], does not even make it to the pending folder. Until a human fills in the owner, it is not treated as a decision. This is the structural guard against "we held the meeting, but nothing moved."

Decisions that pass are turned into pending atoms by promote.py, and once a human approves them at the weekly review gate, they are promoted to official atoms. The principle applied here is the decision_summary_not_clickup_mirror atom (§17.1.2). The task board tracks "what to do"; the decision database tracks "why we decided it." Mix the two and both break.

These three scripts are the skeleton, and AI is the assistant that fills the blanks inside it. Reverse the order — let AI build the skeleton — and hallucinations destroy the very credibility of the decision database.


17.4.3 The Most Dangerous Spot — Why Position 2 Comes Last

Position 2 (STT transcript → auto-generated meeting-notes draft) is the spot every team wants to do first. The picture — "just toss in the recording and meeting notes come out" — is simply too attractive. And that very attraction is why it fails most expensively.

The failure comes in four shapes.

flowchart TD
    R[STT transcript → AI draft] --> R1["Hallucinated decisions
Records as decided what was never agreed"] R --> R2["Missed decisions
Lets actual decisions slip through"] R --> R3["Speaker misattribution
Swaps who proposed what"] R --> R4["Tone flattening
Dissent and nuance lost"] R1 --> X["Collapse of decision tracking —
the reason meeting notes exist"] R2 --> X R3 --> X R4 --> X style X fill:#ffd9d9,stroke:#c0392b

The deadliest of these is the hallucinated decision. Someone in the meeting merely floated an opinion — "Wouldn't 0.5 seconds be better for the global cooldown (GCD)?" — and the AI draft writes "agreed: global cooldown 0.5 seconds." Three weeks later that one line is reflected in the data sheet, the combo design gets built on top of it, and QA cases get written. A decision that was never agreed on propagates irreversibly.

So Position 2 runs under absolute rules.

This does not mean "never do Position 2." Once Positions 3, 4, and 1 are stable and the facilitator has learned the limits of AI output firsthand, Position 2 is well worth adopting. The point is that it comes last.


17.4.4 Position 3 — Why Decision Slot Enrichment Is the Top ROI

This is where six months of operation produced the biggest payoff. A human declares that the decision exists, and AI fills in the decision's supporting fields. The decisive difference from Position 2 is that a human pins down the very fact that there is a decision, first.

AI drafts the three things that take too much time to fill in by hand:

The crux is a prompt that forces evidence quotation and explicitly bans hallucination. Below is the full text of the enrichment prompt I actually ran.

Worked Transcript — Decision Slot Enrichment

Prompt (human input):

Look at the meeting transcript below and fill in only the four fields —
owner / rationale / affected_atoms / follow_up — for decision D1
("unify the combat global cooldown at 0.5 seconds"), which the facilitator declared.
Do not create new decisions. In the rationale, quote transcript statements verbatim
with their timestamps, and for any field with no evidence, do not guess —
write "no evidence."
// (Intent: affected_atoms only from systems/sheets mentioned in the transcript, max 3)

[Meeting transcript]
(14:18) teammate_a: The GCD is 0.3 seconds right now, and inputs keep getting eaten when I chain skills.
(14:22) teammate_b: In yesterday's combo test too, at 0.3 seconds the second skill failed to come out
        in about 30 percent of cases. I kept the logs.
(14:25) teammate_a: Then let's raise it to 0.5 seconds. I'll reflect it in the combo sheet.
(14:26) Facilitator: We go with 0.5 seconds. Decided.
(14:31) teammate_c: If healing skills are on the GCD, you can't use them in an emergency — what about...
(14:33) Facilitator: Let's take that up separately next time.

Claude output (raw, as is):

1. owner: teammate_a
   Evidence: (14:25) "I'll reflect it in the combo sheet." — teammate_a
   explicitly took ownership.

2. rationale:
   - (14:18) teammate_a: "inputs getting eaten when chaining skills at GCD 0.3 seconds"
   - (14:22) teammate_b: "in the combo test, at 0.3 seconds the second skill failed
     to fire in about 30% of cases, logs kept"
   → The dropped-input problem at 0.3 seconds is the direct basis for the 0.5-second decision.

3. affected_atoms (candidates based on transcript mentions):
   - atoms/combat/global_cooldown
   - combo design sheet (transcript 14:25 "combo sheet" mention)
   - atoms/combat/skill_recovery  ← note: at (14:31) the healing-skill
     GCD exemption discussion was "deferred to later" (14:33). Unclear whether it
     falls within this decision's scope. Weak evidence; needs human confirmation.

4. follow_up:
   - teammate_a: reflect GCD 0.5 seconds in the combo design sheet
   - [Separate agenda item] Whether healing skills are exempt from the GCD is
     outside this decision's scope; split it into the next meeting's agenda
     (facilitator's statement at 14:33)

Human verification/rejection:

The facilitator reviewed the output. The owner and rationale quotes were accurate, so they were adopted as is. For the third affected_atoms candidate, skill_recovery, the AI itself had flagged "weak evidence; needs human confirmation," and the facilitator's judgment was to exclude it from this decision's scope — the healing-skill exemption is a matter for a separate decision, not an effect of this D1. The follow_up suggestion to "split it off as a separate agenda item" was adopted and registered on the next meeting's agenda.

What matters here is that the AI did not bulldoze the uncertain item through as a hallucination but reported its own uncertainty. The prompt's constraint — "no guessing, no hallucination; if there is no evidence, state 'no evidence'" — is what produced this honest output. Remove the constraint and the AI confidently puts skill_recovery into affected_atoms, and that hallucination propagates into the graph.

The reviewed decision block passes decision_parser.py — all four fields filled, so no [MISSING] — and moves on as a pending atom.


17.4.5 Position 4 — Atom Routing and Position 5 — Relationship Extraction

When a pending atom that cleared Position 3 is promoted into an official folder, AI recommends which folder to send it to (Position 4).

For this atom ("unify combat global cooldown at 0.5 seconds / owner teammate_a"),
pick up to 3 of the folders below, in priority order, as the best place to file it.
Do not propose creating new folders — choose only from this list.
- atoms/combat/  atoms/character/  atoms/operations/  atoms/visual/

"No creating new folders" is the key constraint. Drop it and the AI proposes folders like atoms/combat_timing/ and atoms/gcd_rules/ without end; categories multiply unboundedly, and search and auto-injection collapse. The principle is to keep categories small and orthogonal, unchanged for a year or more. AI picks only from within that closed list.

Position 5 (relationship extraction between atoms) is adopted last and most cautiously. This is the spot that infers dependency relationships among promoted atoms.

New atom A: "Healing skills are exempt from the global cooldown"
Existing atom B: "Global cooldown unified at 0.5 seconds"

Inferred relationships:
  A.exception_of: [B]
  A.derives_from: [B]
  B.affects: [A]   ← reverse direction assigned automatically

The problem is that this inference is directly exposed to LLM non-determinism. The same input produces different relationships yesterday and today. There are three mitigations — temperature=0 plus a fixed seed on models that support it; a review gate where candidates are proposed and a human approves; and extracting only one direction, with a script deterministically filling in the reverse. Hand both directions to the LLM and one side goes missing.


17.4.6 One or Two at a Time — Rollout Order Is the Safety Mechanism

Turning on all five points at once is the most common and most expensive failure. The operating burden arrives before the benefits do, and the team abandons the system wholesale. Here is the order I actually followed.

Step 1 · Position 3, decision slot enrichment 1–2 months · Start with the No. 1 ROI, lowest risk Step 2 · Position 4, atom routing recommendations +1 month · Enforce a closed list of folder candidates Step 3 · Position 1, STT Once self-hosted infrastructure is in place · Avoid external APIs for security Step 4 · Position 2, meeting-notes draft After the three above are stable; the most cautious · Never auto-commit Step 5 · Position 5, relationship extraction

Putting Position 2 last is the heart of this order. You do the spot you most want to do, last — counterintuitive, but staffing the most dangerous inspection station with the most practiced hands is a basic workshop safety principle.

The order is rational on cost as well. At 100 meetings per month, Position 3 runs about $5–10 and Position 4 about $1–2 (author's estimate from my own operating environment, unverified), so turning on just those two stays under $10 a month. The two highest-impact spots are the cheapest.


17.4.7 Before / After — One Meeting, Two Sets of Notes

The difference between recording the same meeting in two ways is the summary of this whole chapter.

Before — free-form meeting notes (no AI, or Position 2 trusted with the decision slots too):

## 2026-06-06 Combat Sync Meeting

Discussed GCD. Opinions that 0.3 seconds is too short.
Apparently there were problems in the combo test. 0.5 seconds came up.
Healing-skill exemption also briefly mentioned.
Overall the mood seemed to settle toward 0.5 seconds.

Open these notes again three weeks later, and nobody can reconstruct whether "the mood settled toward 0.5 seconds" was a decision or an opinion, who agreed to put it in the sheet, or whether the healing-skill exemption was decided or deferred. No speakers, no owner, and the evidence is "somewhere in the transcript," so you are back to listening to the recording.

After — decision slots + Position 3 enrichment:

## 2026-06-06 Combat Sync Meeting

### Agenda Summary (AI-assisted)
- Dropped-input problem with the 0.3-second combat global cooldown (GCD)
- Whether healing skills are exempt from the GCD (split into a separate agenda item)

### Decisions  (human-declared + AI-enriched)
D1:
  decision: Unify the combat global cooldown at 0.5 seconds
  owner: teammate_a
  rationale: |
    - (14:18) teammate_a: inputs getting eaten when chaining skills at 0.3 seconds
    - (14:22) teammate_b: combo test at 0.3 seconds, second skill failed to fire ~30%, logs kept
  follow_up: teammate_a — reflect GCD 0.5 seconds in the combo design sheet (by 6/13)
  affected_atoms: [atoms/combat/global_cooldown, combo design sheet]

### Split-Off Agenda Items
- Healing-skill GCD exemption → next meeting (facilitator decision at 14:33)

Three weeks later, these notes have been read by decision_parser.py and wired into the graph, and anyone asking "why 0.5 seconds?" gets an immediate answer from the two quoted lines in the rationale. The owner is explicit, so whether the follow_up was executed can be tracked, and even the fact that the healing-skill exemption is a deferred agenda item, not a decision is preserved.

What made the difference is not the amount of AI but a structure that preserves the human's place to declare decisions while assigning AI only the filling-in of evidence. Delete every paragraph the AI filled in the After notes (the agenda summary, the rationale quotes, the affected_atoms candidates), and what remains is a single decision line and an owner — more than half the information in the notes came from AI enrichment, but the crux is that all of that half is evidence quotation that passed human review.


Key Takeaways


Beyond Games. The principle "a human declares that a decision exists; AI fills in only the evidence, the owner, and the impact" is a safety line that applies, exactly as is, to any office worker using AI to write up recordings — not just in games. The most tempting spot (auto-generating full meeting notes straight from a recording) is the most dangerous because of the hallucination that turns the opinion "wouldn't 0.5 seconds be better?" into the decision "agreed on 0.5 seconds." For example, when an HR team writes up a performance-review meeting, have the facilitator personally pin down only the decision — "finalized at grade B" — and ask the AI only this: "Quote the statements in the recording that support this grade; if there are none, say so." Let AI create the decisions, and an evaluation that was never agreed on ends up in someone's HR record, irreversibly.


Try It Yourself

setup 1. In your standard meeting-notes format, create a ## Decisions block and enforce the four fields decision / owner / rationale / follow_up on every decision. 2. Write decision_parser.py — if any of the four fields is empty, print [MISSING <field>]; in particular, if owner is empty, block promotion. 3. Write down the rule that the decision summary is an independent asset holding the "why," not a mirror of the task board (decision_summary_not_clickup_mirror).

prompt 4. Use the Position 3 enrichment prompt. Constraints it must include: "do not create new decisions / quote transcript evidence with timestamps / if there is no evidence, write 'no evidence' / no guessing, no hallucination." Request the four slots: rationale, owner, affected_atoms, follow_up. 5. In the Position 4 routing prompt, include "no creating folders that are not on the list + the closed folder list."

verify 6. Run the enriched decision block through decision_parser.py and confirm there is no [MISSING]. 7. Have a human directly review every affected_atoms entry the AI flagged as "weak evidence," excluding or adopting each one. Never auto-commit, under any circumstances.

Solo Scale-Down If you work alone or have no time to install tooling, skip the scripts and just hand-write a four-line decision block (decision / owner / rationale / next action) at the end of your meeting notes. Write the name even when the owner is yourself. Ask the AI only this: "Quote the evidence for this decision from my meeting memo; if there is none, say so." Even without a pipeline, a place where decisions are declared and a prompt that forces evidence quotation — those two alone are enough for meeting notes to start becoming a decision database.

Part 18 · Decision Impact

18.1 The Decision-Tracking System

We were in the middle of a quarterly meeting. The combat designer proposed unifying the global cooldown at 0.5 seconds, and everyone nodded. Then the senior designer in the next seat raised a hand. "Doesn't this conflict with the 0.3-second decision we made in Q4 last year? Why did we go with 0.3 back then?" The room went quiet for a moment. Nobody remembered the rationale for that decision. We dug through the meeting minutes and found a single line: "Discussed in the combat task force." We ended up spending 30 minutes reconstructing last year's decision, and even then we never found out why it was 0.3.

Decisions are harder to track than to make. When hundreds pile up in a year, no human head can keep up with which decisions are alive, which have been retired, and which rest on other decisions as premises. This chapter covers a system that pins decisions down as atoms and turns them into a trackable asset. The core is simple. Record each decision as a card with decision_id, owner, and rationale; connect the cards with wikilinks to form a graph; and trace backward with grep to see how far the impact spreads.

18.1.1 The Decision Card: Pinning Decisions Down as Atoms

The minimum unit of decision tracking is the decision card. I brought over, as is, an actual card from Project A (an MMORPG in development) that I run. It is the very 0.5-second unification decision that collided in the meeting above.

---
decision_id: D2026_Q2_017
title: Unify the combat global cooldown at 0.5 seconds
type: system_change
status: active        # active / superseded / deprecated
created: 2026-04-18
owner: teammate_a      # combat designer; proposed and owns the decision
approved_by: Minsoo Lee    # Design Director
approval_meeting: 95_BattleTF_2026-04-18

scope:
  - combat_system
  - all_active_skills

content: |
  Apply a 0.5-second global cooldown to all combat active skills.
  Healing skills are an exception (separate decision D2026_Q2_018).

rationale:
  - Combo input readability problem (accumulated player feedback)
  - Simulations point toward longer average combat length
  - Flatten the learning curve for new players

affected_atoms:
  - combat_global_cooldown_constant
  - combat_skill_cooldown_rule

affected_files:
  - CombatBalance.xlsx
  - CombatFormula_v3.md
  - UI/skill_cooldown_indicator

implementation:
  target_build: 2026-05-09
  impl_owner: teammate_b    # code lead
  qa_owner: teammate_c      # QA senior

related_decisions:
  - supersedes: D2025_Q4_034   # the previous 0.3-second decision
  - relates_to: D2026_Q2_018   # healing exception
---

Three fields are the spine. decision_id gives the decision a permanent address. owner nails down who is responsible for this decision. rationale answers the "why did we do that?" six months later. The "why was it 0.3?" the meeting failed to find is exactly what should have been sitting in the rationale field of D2025_Q4_034. The remaining fields (scope, affected_atoms, related_decisions) are wiring for impact tracing and graph connections.

One piece of deliberate design goes in here. Force all 12 fields and people start dodging card writing altogether. So I split them into 5 required fields (decision_id, title, owner, status, rationale) and 7 optional ones. Fill in just the 5 required fields right after the decision is made in the meeting and the card is alive; the rest get filled in during implementation.

18.1.2 The Full Decision-Tracking Flow

The skeleton of the tracking system is the path a single card travels from creation to retirement. Pay attention to where the irreversible gates sit.

flowchart TD
    A[Decision arises in a meeting or team messenger] --> B[Draft the decision card
5 required fields] B --> C[Assign decision_id and register in the index] C --> D{Impact scope analysis
impact} D --> E[Fill in affected_atoms and affected_files] E --> F[Connect into the graph via wikilinks] F --> G{owner and approved_by review gate} G -->|Rejected| B G -->|Approved| H[Applied to the build] H -.irreversible.-> I[Propagates to other docs and decisions] I -.irreversible.-> J[Post-hoc measurement and verification] J --> K{Evolution check} K -->|Superseded| L[status: superseded
supersedes link] K -->|Still valid| M[status: active retained] style H fill:#ffe0e0 style I fill:#ffe0e0

Everything from drafting (B) to the review gate (G) is reversible. Fixing or discarding the card costs almost nothing. After the change lands in the build (H), though, it is effectively irreversible. Even if a hotfix rolls back a change players have already felt, it leaves a mark on community perception, and once follow-up decisions start stacking on top of this one as a premise, the cost of reversal grows exponentially. So every review by the decision-maker has to finish at gate G. This is exactly the same structure as the "recording and casting are irreversible stages" principle covered in Part 5.

18.1.3 The Decision Graph: Connecting the Cards

Once cards are atoms, they can be connected to each other. supersedes and relates_to in related_decisions become the edges of the graph. The collision in that meeting was, in fact, one small piece of this graph.

D2025_Q4_034 Global cooldown 0.3s (deprecated) D2026_Q2_017 Global cooldown 0.5s (active) D2026_Q2_018 Healing-skill cooldown exception (active) D2026_Q2_025 PvP global cooldown variant (active) supersedes relates_to relates_to

With this graph, the meeting would have ended in 30 seconds. Open D2026_Q2_017 and supersedes: D2025_Q4_034 is right there; one click into that card's rationale and the "why 0.3" is laid out as is. The graph is the evolution history of the decisions, and the evolution history of the decisions is the history of the game. Even branches derived from the main decision, like the PvP variant (D2026_Q2_025), are traceable at a glance.

18.1.4 Extracting the Impact Scope Automatically — impact

If a human fills in a decision card's affected_atoms and affected_files one by one, things get missed. Project A has an impact-scope extraction procedure called impact. It takes a decision atom and sweeps the graph in three directions.

The union of the three paths is the decision's true impact scope. On top of that, the portal_layer_change_impact_check atom separately checks whether the decision touches the portal layer (externally exposed documents and API specs). If the portal layer is involved, the grade goes up one level — external propagation is more expensive to undo.

18.1.5 Worked Transcript: From Meeting Notes to Decision Card

That is it for theory. Here is the full process of actually throwing a chunk of meeting notes at an LLM and getting a decision card back, with the prompts in full and the raw output included as is. No summarizing. You see where Claude got confused, where the human rejected, and the rework request — all of it.

First Prompt (in Full)

Turn the meeting notes below into a decision card in YAML. The schema is
decision_id / title / type / status / owner / approved_by / scope /
content / rationale / affected_atoms / affected_files / related_decisions.
owner is the proposer; approved_by is the approver. For affected_atoms and
related_decisions, use only what actually appears in the notes — do not guess.
// (intent: decision_id is D{YYYY}_Q{n}_{seq}; if seq is unknown, XXX)

[Meeting notes]
- Decided to unify the combat global cooldown at 0.5 seconds (proposed by teammate_a).
  Cooldowns currently vary per skill, and feedback keeps piling up that combo readability is poor.
- Last year's 0.3 seconds was called out as too fast — combos got mushed together.
- Healing skills can't handle crisis response at 0.5 seconds. Needs an exception.
- Approved by the director (Minsoo Lee). teammate_b implements, teammate_c handles QA.
- Target build is May 9.

Claude's First Output (Raw, Unedited)

---
decision_id: D2026_Q2_XXX
title: Unify the combat global cooldown at 0.5 seconds
type: system_change
status: active
owner: teammate_a
approved_by: Minsoo Lee
approval_meeting: 95_BattleTF
scope:
  - combat_system
  - all_skills          # ← healing skills are an exception, yet it grabbed all_skills
content: |
  Apply a 0.5-second global cooldown to all combat skills.
rationale:
  - Accumulated feedback on combo readability
  - Last year's 0.3 seconds was too fast; combos got mushed
affected_atoms:
  - combat_global_cooldown      # ← differs from the real atom name. A guess
affected_files: []              # ← left empty. Not specified in the notes
implementation:
  target_build: 2026-05-09
  impl_owner: teammate_b
  qa_owner: teammate_c
related_decisions:
  - supersedes: D2025_Q4_XXX    # last year's 0.3-second decision. seq unknown
  - relates_to: healing-skill exception decision (ID undetermined)  # ← filled with a description, not a link
---

Honestly assessed, this output scores about 60 out of 100. The good and the wrong are mixed together.

Human Verification and Rejection

I did not accept this card as is. I wrote up (1), (2), and (3) as rejection reasons and requested a rework. The heart of the rejection is the guessed atom name. If an atom ID is off by even one character, every grep and backreference downstream misses. The accuracy of the tracking system is bound to the accuracy of its atom IDs.

Second Prompt (the Rework Request, in Full)

Fix three things. For scope, healing skills are an exception, so use
all_active_skills instead of all_skills, and put the healing exception in
content. For affected_atoms, I told you not to guess and you guessed —
if you don't know the real atom name, leave [needs verification] and only
add a comment on what to check; I'll check the index myself. For
related_decisions, set the healing exception to relates_to D2026_Q2_018,
and supersedes to D2025_Q4_034. affected_files as [needs verification]
candidates only.

Claude's Second Output (Raw)

affected_atoms:
  # [needs verification] Candidates below. Confirm exact atom IDs against the index.
  - combat_global_cooldown_constant   # global cooldown constant value
  - combat_skill_cooldown_rule        # per-skill cooldown rule (includes exception handling)
affected_files:
  # [needs verification] Not in the notes. Usual impact candidates for type=system_change:
  - CombatBalance.xlsx        # cooldown values sheet
  - CombatFormula_v3.md       # combat formula doc
  - UI/skill_cooldown_indicator   # cooldown UI indicator
related_decisions:
  - supersedes: D2025_Q4_034
  - relates_to: D2026_Q2_018
scope:
  - combat_system
  - all_active_skills
content: |
  Apply a 0.5-second global cooldown to all combat active skills.
  Healing skills are an exception (separate decision D2026_Q2_018).

The second pass came back honest. Instead of guessing atoms and asserting them, it attached [needs verification] flags with supporting comments. I opened the atom index, confirmed that the two names combat_global_cooldown_constant and combat_skill_cooldown_rule actually exist, and removed the flags. The three affected_files candidates were also confirmed against the index. The final card at the top of this chapter is the result.

The lesson of this transcript is a single one. An LLM is powerful as the drafter of decision cards, but the final confirmation of atom IDs and decision IDs has to be done by a human checking against the index. The AI explores candidates; the human adopts them. Mix the two roles, and a wrong atom name contaminates the entire graph.

18.1.6 Tracing Impact Backward with grep

Once the cards and the graph are bound together by atom IDs, "where does this decision have impact" is answered with one line of grep. Take decision D2026_Q2_017's core atom, combat_global_cooldown_constant, and sweep the manuscripts, sheets, and all decision cards for backreferences.

rg "combat_global_cooldown_constant" --type md --type yaml -l
# → D2026_Q2_017.yaml          (the decision card itself)
#   D2026_Q2_025.yaml          (PvP variant — re-cites this constant)
#   CombatFormula_v3.md        (formula doc)
#   95_BattleTF_2026-04-18.md  (original meeting notes)

This result is the impact map: change this constant and four places shake. The fact that the PvP variant card cites the same constant is easy for human memory to miss, but grep does not miss it. It works because the atom ID was exact — grep for the first output's combat_global_cooldown and not one of these four lines would have matched. Grade classification (§18.2), the full-cycle workflow (§18.3), and the refined grep workflow (§18.4) all stand on this atom ID accuracy.

18.1.7 The Difference a Tracking System Makes

Here is a before-and-after comparison from my Project A. The figures below are the author's estimate (unverified); read them for direction and ratio rather than absolute values.

Item Without the system With the system Direction
"Did we decide this before?" re-discussions 8–12 per quarter 0–2 per quarter Sharply down
Mapping a decision's impact scope 1–2 days Minutes with grep Sharply shortened
Tracing decision evolution history Reliant on senior memory Automatic via the graph Human dependency removed
New team members learning decision history 1–2 months 1–2 weeks Biggest effect

The biggest effect is the last row. Instead of cornering a senior to ask "why is this game the way it is," a new team member reads down the decision graph on their own. Decision tracking becomes the company's decision-making learning asset. That said, the first quarter after the system comes in does carry a real card-writing burden. The safe path is to establish the 5 required fields first and expand gradually.

18.1.8 From Conservative to Progressive — Automation Stands on Atom Decomposition

The operation so far is the conservative application. Humans decide in meetings, write the cards, and identify the affected atoms; automation handles only indexing, search, grep, and graph visualization. Humans own the core judgment; automation owns storage and retrieval.

The next step is the direction the transcript above pointed to. With raw meeting notes as input, an LLM drafts all 12 fields of the decision card, explores affected-atom candidates along the graph, and even recommends a grade. What remains in human hands narrows to two things: reviewing whether the AI-filled card and atom names match the index, and final approval. The burden of filling 12 fields from zero and the burden of checking an LLM draft's atom names against the index are different in kind.

For this progressive application to take root, three skeletons are needed. First, a decision graph in which every decision is registered as an atom and connected by wikilinks. A blob of meeting notes cannot serve as automation's input — it has to be decomposed into decision units. Second, an automatic impact grader (§18.2) that computes the number of affected domains, the rollback cost, and the player impact scope on top of the graph and recommends a grade. Third, grep and LLM impact tracing (§18.4) that operates precisely on atom IDs and wikilinks.

Here the message that runs through this whole book surfaces once more. Decomposing decisions into atoms, graphs, and grades has "convenience of search and backreference" as its surface, but the essence lies in the fact that from an undecomposed blob of meeting notes, automated impact analysis cannot even tell what the unit of a decision is. The general thesis (§6.6) — decomposition holds unifying the collaboration language as its surface purpose and the precondition for procedural automation as its essential purpose — appears in the decision-making domain as the decision graph, atoms, and grades. It is the same skeleton as the world BT (Behavior Tree) and quest cloud in Part 5 and progressive balancing in Part 8. The theory was possible in the 2010s too, but automatically decomposing meeting notes into decision atoms was the blocker; after 2023, with LLMs taking on the draft of that decomposition, much of the vision that lived only on paper entered the realm of the practicable.

Key Takeaways

Beyond Games. The decision card is a device that lets any organization — not just a game team — answer "why did we decide it that way back then," even six months later. The marketing team wasting 30 minutes because "we decided to drop this channel last quarter — why was that again?" cannot be found in a one-line meeting note: that disappears with a single three-field card of decision_id, owner, and rationale. For example, when HR sets a policy like "standardize on two remote days per week," write the proposer, the approver, the rationale (productivity data, employee survey), and the ID of the superseded previous policy on that one card, and a year later, at the policy review, the grounds for the past judgment are still alive.

Try It Yourself

The minimal web-chatbot path (no terminal) — The core of this chapter is not the decision-card directory or grep; it is the idea of pinning a permanent address (decision_id), an owner (owner), and a rationale (rationale) onto each decision, and looking up past decisions before making a new one. That idea reproduces with nothing but a web chatbot (ChatGPT or Claude on the web), no CLI or atom index required. The three steps below are the main road. 1. Write one decision as one line. A single ordinary document named decisions.md is enough. No YAML, no scripts. - [D17] Unify global cooldown at 0.5s (owner: me, rationale: combo readability, supersedes: D08) 2. When turning meeting notes into cards, paste the following into the web chatbot. It carries over the four constraints of the first prompt as is. Turn the decisions in the meeting notes below into a table. Columns: decision_id / title / owner / rationale / superseded past decision. If you can't identify the owner, put [MISSING]; if you don't know atom or file names, put [needs verification]. Do not guess. // (intent: decision_id is D{year}_{seq}; if seq is unknown, XXX) [meeting notes body] 3. Before making a new decision, search decisions.md first with in-document find (Ctrl+F) — that alone settles the one question, "did we decide this before?" This is the hand-cranked version of grep backtracing. Bring in the atom index, YAML cards, and the rg workflow only when decisions number in the hundreds and a single document becomes too heavy to search.

setup (the infrastructure version — once the minimal path above feels routine) — Create a decision-card directory and an index file.

decisions/
  D2026_Q2_017.yaml
  _index.json        # by_status / by_scope / by_quarter aggregates

prompt — When you throw a meeting's decision items at an LLM, always include the four constraints from the first prompt above. In particular, state "do not guess atom names; leave them as [needs verification]."

verify — Check the produced card's affected_atoms entries against the atom index, confirm the real names, then remove the flags. Then run rg "<atom_id>" -l on the core atom and cross-check that the impacted files match the card's affected_files.

Solo Scale-Down

If you are working alone with no team infrastructure, drop the YAML card. Write one decision as one line of Markdown.

- [D17] Unify global cooldown at 0.5s (owner: me, rationale: combo readability, supersedes: D08's 0.3s)

Stack these one-liners in a single decisions.md file, and before making a new decision, search past decisions first with rg "cooldown" decisions.md. No cards, no graph, no tooling — but the one question, "did we decide this before?", is solved. Ninety percent of a tracking system starts with this one-line habit.

18.2 Impact Propagation and Tier Classification

I was tidying up the meeting notes after a meeting. One decision sat there in a single line: "Unify the global cooldown at 0.5 seconds." In the meeting, agreement took less than 30 seconds. Everyone nodded, and we moved on to the next item.

That one line ate the next two months. All 277 skills in the combat data were affected, the UI's cooldown gauge animation had to be redrawn, and the balance spreadsheet was torn up and rebuilt twice. Another decision written in the same meeting notes — "fix typos in the tutorial guidance text" — was done in five minutes.

On the page, both decisions were equally one line. Even their character counts were similar. Yet one took five minutes and the other took two months. Making that difference visible at the moment the meeting notes are written — that is impact tier classification. When the tier is invisible, a two-month decision gets buried in the same line as a five-minute one.

This chapter covers how to automatically classify a decision's impact into five tiers, and how to trace how far that impact spreads across the decision atom graph. The tools are the decision atoms and the impact extraction built in the previous chapter, plus the portal_layer_change_impact_check atom.


What Happens When Tiers Are Invisible

First, let me describe what things look like without tier classification. When every decision sits on the same line, two kinds of accidents take turns blowing up.

One is under-handling. A decision that shakes the whole quarter, like the global cooldown one, gets treated as a "five-minute" item and enters the build without verification. The impact only surfaces two months later, and by then the cost of rolling it back has piled up like a mountain.

The other is over-handling. Fixing a single typo convenes a task force (TF) and requires the game director's sign-off. Decision cycles explode, and the time the director should be spending on T0 decisions gets sucked into typo meetings.

The two accidents look like opposites, but they share the same root: the weight of a decision is invisible. Because the weight is invisible, we spend effort on light decisions and let heavy ones slip through. Tier classification is the work of attaching a weight label to each decision, and the moment the label is attached, the handling path splits automatically.


18.2.1 The Five Impact Tiers — T0 Through T4

On Project A — the MMORPG project I run — decision impact is divided into five tiers. The higher the tier, the heavier the decision, and the more people and time it takes to handle.

Tier Definition Examples Decider Cycle
T0 Game vision, core systems Mobile-first decision, core mechanic change Game director + CEO Quarterly
T1 Systems, multi-discipline Global cooldown unification, adding a new class TF chair + director 1–2 weeks
T2 Single discipline, mid-size Adjusting a specific skill's numbers, adding a UI component Discipline director 3–5 days
T3 One-off, small Editing a single NPC's dialogue, fine-tuning a color One senior 1–2 days
T4 Immediate, hotfix Bug fix, text typo Owner Hours

The table looks textbook-clean. But the practical difficulty is not memorizing the table — it is judging which cell the decision in front of you goes into. You need to know that "global cooldown unification" is T1 at the moment you write the meeting notes, not after the meeting is over. That is why the three criteria in the next section are the core.


18.2.2 The Three Criteria That Decide the Tier

Tiers are not assigned by gut feeling. We evaluate three criteria and adopt the highest tier among them.

Tier Decision Matrix — 3 Criteria × 5 Tiers Criterion \ Tier T0 T1 T2 T3 T4 Affected disciplines 5+ 2–4 1 1 1 Rollback cost Very high High Medium Low Very low User impact scope All High Medium Low Very low

Of the three criteria, the number of affected disciplines can be counted mechanically on the decision atom graph. Collect the discipline tags (combat, UI, data, narrative, and so on) of the atoms the decision touches, and you are done.

The problem is the other two. Rollback cost and user impact scope do not reduce to numbers on the graph. "How much would it cost to undo this decision two months from now" is a natural-language judgment. This was the last wall standing in front of automatic impact classification before 2023. The discipline count had been automated, but the two natural-language cells stayed empty, so in the end a person regraded everything from scratch anyway. The wall came down when LLMs started reading the body of a decision atom and drafting those two cells.

Let me be honest here. What the LLM fills in is a draft, not a verdict. Even if the LLM estimates the rollback cost as "high," a discipline director can rule, "Given our sheet structure, this is medium." Automatic classification does not replace human judgment — it keeps the human from starting with empty cells.


18.2.3 Worked Transcript — Having the LLM Classify One Decision

This is the actual process of feeding one decision atom from the previous chapter into the LLM, as is, and having it assign a tier. I am reproducing the whole exchange without summarizing, including the rejection and the follow-up request.

Input — The Decision Atom

# decisions/D2026_Q2_017.md (the atom registered in the previous chapter, 18.1)
id: D2026_Q2_017
title: Unify global cooldown at 0.5 seconds
date: 2026-04-18
proposer: teammate_a
decider: Minsoo Lee    # Design Director (same as approved_by in 18.1)
context: Per-class cooldown baselines vary, so balance recalculation costs
         pile up every time a new class is added. Proposal to unify on a 0.5-second global baseline.
affected_atoms: [combat_cooldown_rule, skill_table, ui_cooldown_gauge]
reversible: ?
status: proposed

The Full Prompt

Assign an impact tier from T0–T4 to the decision atom below.
Evaluate the three criteria separately — affected discipline count / rollback cost /
user impact scope — and adopt the highest tier among them. For any cell
you are not confident about, write "uncertain" and say a human needs to review it.
// (intent: T0 = 5+ disciplines, very high cost, all users; T1 = 2–4, high, high; T2 = 1, medium; T3 = 1, low; T4 = 1, very low)

[atom text attached]

Claude's Raw Output (Unedited)

Evaluating the three criteria:

1. Affected discipline count
   From affected_atoms:
   - combat_cooldown_rule → combat
   - skill_table → data
   - ui_cooldown_gauge → UI
   Three disciplines are explicit.
   However, given the context — "balance recalculation every time
   a new class is added" — the balance discipline is also likely
   to be indirectly affected.
   Direct 3 + indirect 1 = effectively in the 2–4 range. → T1 signal

2. Rollback cost
   Undoing the global cooldown unification would mean re-splitting
   every skill cooldown in skill_table back to per-class baselines.
   The exact cost is uncertain since the row count of skill_table
   is unknown, but "unify → restore the split" is generally
   expensive. → high (T1 signal)
   ※ This cell cannot be finalized without the size of skill_table. Marked uncertain.

3. User impact scope
   Cooldowns apply to every combat situation, so nearly all users
   will feel this. → high (T1 signal)

Adopted tier: T1
Rationale: all three criteria point to T1. It is not T0 —
this does not rise to a "core mechanic change"; it is a parameter
unification within an existing mechanic.

Needs human confirmation:
- Rollback cost cell: actual row count of skill_table required
- The reversible slot is empty ("?") → this classification assumes
  reversible will be filled in as "hard".

Human Verification and Rejection

I did not accept this as is. Item 1 of the output needed a correction. The LLM read the balance discipline as "indirectly affected," but in Project A's actual structure, unifying the cooldown overhauls the balance spreadsheet directly — the balance formula for new classes takes the cooldown value as an input. So balance is not indirect; it is the fourth directly affected discipline. The affected discipline count is 4. Even so, the adopted tier, T1, does not change.

One more thing: the "uncertain" mark on the rollback cost cell was accurate. skill_table actually held 277 skill rows, and undoing the unification would mean re-splitting the cooldown of all 277 rows. The cost was closer to "very high" than "high." Still, one cell moving up does not push the adopted tier to T0 — the discipline count and the user impact remain in T1 range, so the highest tier is still T1.

Follow-Up Request

Correction: add balance_sheet to affected_atoms (the balance formula takes
the cooldown value as an input, so the impact is direct). That makes 4 affected
disciplines, and skill_table has 277 rows.
Apply this, fill in reversible, and finalize the tier.
Applied.

Affected disciplines: combat, data, UI, balance = 4
Rollback cost: very high (re-splitting all 277 rows of skill_table)
User impact: high

reversible: hard (close to effectively irreversible — re-splitting
277 rows requires 1–2 builds plus a quarter of balance re-verification)

Final tier: T1
Caveat: looking at the rollback cost cell alone, it is a T0 signal,
but 4 affected disciplines and "high" user impact are T1 range.
Two of the three criteria point to T1, so T1 stands.
I do recommend marking this decision as "on the heavy end of T1" —
because its rollback cost is T0-grade.

That last sentence was more careful than the human. The tier is T1, but it proposed labeling the decision "on the heavy end of T1" because the rollback cost is T0-grade — a judgment that was nowhere in the original prompt. I took it. I added a weight: heavy tag to this decision and raised its post-tracking priority.

This transcript shows one thing above all. The LLM produces the draft classification and its reasoning; the human corrects it with domain facts (balance is directly affected; 277 rows). Neither works alone. A human alone starts from empty cells and is slow; an LLM alone writes "high" without knowing about the 277 rows.


18.2.4 Code That Traces Impact Propagation Automatically

Once the tier is set, the next question is "how far does it spread?" We apply the impact extraction from the previous chapter — inbound edges, the ontology's affects relation, wikilink reverse references — to the decision atom.

# impact_propagation.py — trace a decision atom's propagation scope

def trace_impact(decision):
    # 1st hop: atoms and files the decision touches directly
    direct = decision.affected_atoms + decision.affected_files

    # 2nd hop: atoms that reverse-reference 1st-hop atoms via wikilink (impact inbound edges)
    secondary = []
    for atom in direct:
        secondary.extend(find_inbound_refs(atom))   # [[atom]] reverse references
        secondary.extend(find_affects_edges(atom))   # ontology affects

    secondary = dedup(secondary) - set(direct)

    return {
        "direct": direct,
        "secondary": secondary,
        "affected_fields": determine_fields(direct + secondary),
        "estimated_hours": estimate_hours(direct, secondary),
    }

The key is find_inbound_refs — the function that collects the incoming arrows on the atom graph, the ones pointing at a given atom via [[...]]. What a decision touches (outgoing arrows) is written in the atom itself, but who depends on those atoms (incoming arrows) only becomes visible by scanning the whole graph in reverse. Two-month-scale impact almost always hides on this inbound edge side.

Here, honestly, is what running this trace on D2026_Q2_017 produced. direct was the 4 atoms confirmed above. secondary was the atoms reverse-referencing skill_table — skill description text, skill icon mappings, per-class skill trees, and more — which came spilling out in a chain. The numbers vary by point in time, so I will not pin them down. What the trace established is the direction — "secondary is tens of times larger than direct" — and the exact atom count depends on the state of the graph. The direction alone is enough: when secondary is an order of magnitude larger than direct, that is a T1 signal, and it means the decision is a post-tracking target.


18.2.5 Where Tier Classification Plugs In — Mermaid

Classification is not a standalone step; it is fixed as a gate in the middle of the decision flow. When a decision candidate is registered, the automatic analysis recommends a tier, and only after a human reviews and adjusts it does the decision move on to the decision meeting.

flowchart TD
    A[Decision candidate registered
meeting notes → draft decision atom] --> B[Automatic impact analysis
impact extraction] B --> C{Automatic evaluation of 3 criteria} C -->|Affected discipline count| D1[Computed from the graph] C -->|Rollback cost| D2[LLM draft → uncertainty flagged] C -->|User impact| D3[LLM draft → uncertainty flagged] D1 --> E[Adopt the highest tier
T0–T4 recommendation] D2 --> E D3 --> E E --> F{Human review} F -->|Correct with domain facts| G[Tier finalized + weight tag] F -->|Reject and reclassify| C G --> H{Tier branch} H -->|T0| T0[Director + CEO / quarterly cycle] H -->|T1| T1[TF / 1–2 weeks] H -->|T2| T2[Discipline director / 3–5 days] H -->|T3| T3[One senior / 1–2 days] H -->|T4| T4[Owner handles immediately] T0 --> I[Merged into the build — irreversible] T1 --> I T2 --> I T3 --> I T4 --> I I --> J[Post-tracking
estimate vs actual → feeds the next estimate] J -.absorbed as reversible.-> B classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; class B,C,D1,E,I code; class D2,D3 ai; class F,G,T0,T1,T2,T3,T4 human; class A data;

In this flow, there is exactly one irreversible step: merging into the build (I). Everything before it is reversible — if the wrong tier gets recommended, a human can reject it, and a weight tag can simply be removed. It becomes irreversible only after it enters the build and propagates into other documents. That is why the gate (the human review at F) sits in front of the build. A human blocks once, before the irreversible line is crossed.

The final dotted line — the arrow from post-tracking (J) back to the automatic analysis (B) of the next decision — is what turns this system into a learning cycle. Measured data from the irreversible step (for example, QA took longer than estimated) is absorbed into the reversible stage of the next decision.


18.2.6 Post-Tracking — Turning the Gap Between Estimate and Actual into Learning

One week to one month after a decision enters the build, we put the estimate and the actual side by side. This is the post-tracking form for D2026_Q2_017.

Post-tracking for decision D2026_Q2_017  (form example · numbers are fictional inputs)
─────────────────────────────────
Work hours (estimate → actual)
  code:    16h → 22h  (+38%)
  data:     8h →  6h  (-25%)
  UI:       4h →  4h  (=)
  QA:       8h → 12h  (+50%)
  total:   36h → 44h  (+22%)

Affected atoms (estimate → actual)
  direct:    4 →  4   (exact)
  secondary: est. tens → actual tens  (direction matched, exact count not disclosed)

Incidents: 0
Error pattern: QA exceeds the estimate every time (+50% this time)
Apply to next decision: add a default +20% margin to QA estimates

The block above is a form example showing what post-tracking looks like. The hour and percentage values are not real project data — they are fictional inputs filling out the form — so on your own project, fill it in with your own numbers. Exactly as this book promises: we show the structure, and you measure the numbers. One thing is real regardless of the form: the direction of the error — "QA exceeds the estimate every time" — and the procedure of feeding that direction back into the next decision. That is where the prescription comes from: attach a QA margin (say, +20%) to the next estimate from the start.

The value of post-tracking is not in nailing exact numbers but in feeding back the direction of the error. As estimates get more accurate, trust in the tier classification rises; as trust rises, delegation becomes possible.


18.2.7 Accident Patterns and Prescriptions by Tier

Each tier has its own recurring accidents. And its own prescriptions.

Tier Accident pattern Prescription
T0 Vague vision → confusion all quarter Require a one-line vision statement in the decision document
T1 Cross-discipline conflict → schedule slips Every affected discipline sends a representative to the TF
T2 Missed impact on adjacent systems → follow-up decisions explode Mandatory secondary tracing
T3 Small decisions accumulate → consistency erodes Review T3s as a batch in the quarterly retrospective
T4 Insufficient verification → repeat hotfixes Even a hotfix gets at least a one-person review

Every prescription in this table runs on tools from the preceding sections. T2's "mandatory secondary tracing" is find_inbound_refs from §18.2.4, and T1's "every affected discipline sends a representative" is settled by the affected_fields that §18.2.4 captures — that is what determines who needs to be in the room.

The most expensive accident is nowhere in this table: getting the tier itself wrong. When a senior decides a T0 alone, the vision takes damage; when the director personally handles a T4, a bottleneck forms. Get the tier wrong, and every prescription beneath it operates in the wrong place. That is why the human review gate of §18.2.3 is not a mere formality.


18.2.8 Measurement — The Effect of Running Tiers

Here is a before-and-after comparison of introducing tier classification on Project A. The absolute values below are fabricated examples; the direction (the inequality) is the real trend.

Item Without tiers With tiers
Decision cycle Uniform 1–2 weeks for everything Differentiated, from quarterly T0 to hour-scale T4
Mishandled decisions Many per quarter Few per quarter
Director's weekly decision load High (every decision routes to the director) Low (T2–T4 delegated)
Hotfix cycle 1–2 days 4–24 hours
Decision analysis at the quarterly retrospective Hard to group Aggregated as per-tier statistics

The reason the table is written in directions rather than definitive numbers comes down to row 3 alone. Reclaiming the director's time is the biggest effect of tier classification. Without tiers, every decision, from typos to vision, converges on one person — the director. With tiers, T2 and below split off to discipline directors, seniors, and owners, and the director concentrates on T0 and T1. Delegation becoming possible means the director gets back the time for the decisions that are genuinely heavy.


18.2.9 This Chapter's Place in the Progressive-Application Skeleton

The previous chapter built the decision atom graph; this chapter put an automatic tier classifier on top of that graph. They are not tools drifting apart — they are consecutive positions in one skeleton.

flowchart LR
    A["① Decision atom graph
(previous chapter, 18.1)
meeting notes → 12-slot atom"] --> B["② Automatic impact tier classifier
(this chapter, 18.2)
graph → T0–T4 + estimates"] B --> C["③ Wikilink impact tracing
(18.4 grep workflow)
atom ID → propagation scope"] C --> D["Human review and approval
irreversible gate"] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; class C code; class B ai; class D human; class A data;

The three elements run in series. The graph produces the input (①), the classifier weighs it (②), and the impact trace unfolds the propagation scope (③). This chapter is the middle position.

All three entered the realm of the feasible only after LLMs advanced. The last wall to fall was ②'s two natural-language cells — rollback cost and user impact scope (the "last wall" of §18.2.2) — and once LLMs began drafting them, ①→②→③ finally started running in series.

The reversible/irreversible alignment also interlocks with this skeleton. As the §18.2.5 flowchart shows, the only irreversible line is merging into the build; the merge itself cannot be undone, but its measured results come back as reversible learning that makes the next decision more accurate.


18.2.10 Common Failures

Pattern Prescription
Handling every decision on the same cycle Differentiate cycles by tier
No tier classification; everyone judges on their own A three-criteria classification gate
Ignoring tiers (a senior decides a T0) Enforce the decider table
Skipping the pre-decision impact assessment Make automatic analysis a gate before the decision meeting
Finalizing natural-language cells straight from LLM output A human corrects them with domain facts
Closing a decision without post-tracking Compare estimate vs actual at 1 week to 1 month
Not feeding estimate errors into the next decision Apply the error direction to the next estimate's margin

Beyond Games. An impact tier is a label that makes visible, in advance, whether a one-line request is a five-minute job or a two-month one — so it works in any workplace where decisions pour in. When a one-line edit to the company wiki and a "change the leave policy for every department" both arrive as "one agenda item" and ride the same approval line, the light work gets over-handled and the heavy work slips through unverified. For example, when an operations team takes in work requests, attaching T0–T4 by three criteria — number of affected departments, rollback cost, customer impact scope — automatically splits what the owner can handle on the spot from what needs a team lead's sign-off, and the manager's time is reclaimed for the decisions that are genuinely heavy.

Try It Yourself — Classify One Decision

setup. Prepare one decision atom from the previous chapter. Its affected_atoms slot must be filled in. If it is empty, you cannot count the affected disciplines.

prompt. Use the full prompt from §18.2.3 as is. Do not drop the three key lines — (1) evaluate the three criteria separately, (2) adopt the highest tier, (3) mark any unconfident cell "uncertain" and request human confirmation. Without the third line, the LLM will assert things it does not know.

verify. When the LLM output comes back, check two things yourself. First, even if the LLM counted the affected disciplines, recount them directly on the atom graph — it may have promoted an indirect impact to direct, or missed one (the balance case in the transcript). Second, fill the cells marked "uncertain" with domain facts (real scale, like skill_table's 277 rows). Only after both checks do you finalize the tier and pass it to the build gate.

Solo Scale-Down

If you are building a project alone — no team, no TF — five tiers are overkill. Cut it down to three.

Start with zero lines of code on the tooling side, too. When you jot down a decision, just prefix it with a [heavy], [medium], or [immediate] tag. That alone makes you pause once more in front of a decision tagged "heavy" — the essence of tier classification is, in the end, the habit of pausing before heavy decisions, and automation is merely the device that makes that pause keep working at team scale.


Key Takeaways

18.3 Pre- and Post-Decision Impact Tracking Workflow

Three weeks after launch, we sat in a retrospective tracing why PvP balance had collapsed. Working backward across the whiteboard, we arrived at a single decision made a month earlier: "Raise the global cooldown from 0.3 to 0.5." It was a reasonable proposal, agreed on within two hours in response to feedback that combos were hard to read. But that change raised the tank classes' survival rate 14% more than expected, and that broke PvP. Nobody in the decision meeting said the decision would reach all the way to tanks. The decision itself wasn't wrong. The cause of the accident was failing to see how far the decision would spread before deciding.

Impact tracking has to happen in two places: before you commit a decision (pre), to see how far it will spread, and after you apply it (post), to confirm it really spread only that far. This chapter ties those two trackings into a single workflow.


18.3.1 Pre-Tracking and Post-Tracking Read the Same Graph Twice

The core of decision impact analysis is surprisingly simple. Treat a single decision atom as a node, and read the edges coming into that node and the edges going out of it. Pre-tracking asks "if we change this decision, what gets affected?" (outbound plus reverse references); post-tracking asks "did the impact actually land as intended?" (the same edges, compared against measurements).

On my Project A, decisions are kept as atoms in the decisions/ folder. Twenty-six have accumulated so far, and each atom carries the date, the people involved, the rationale, and the impact scope in its frontmatter. The tool that extracts the impact scope is impact, and the atom that enforces those extraction rules at the decision level is portal_layer_change_impact_check. These three are the actual assets behind pre- and post-tracking.

Decision atom D2026_Q2_017 Basis: user feedback Parent decision D_011 Reverse ref: GDD link CombatFormula.md CombatBalance sheet UI combo display Inbound (why this decision) Outbound (where it spreads)

Pre-tracking reads the right side (outbound) to predict "it spreads this far"; post-tracking pulls in measurements from those right-side nodes and compares them against the prediction. The same graph, read twice, for two different purposes.


18.3.2 Pre-Tracking — Reading the Graph Before You Commit the Decision

From here on, this is an actual session. I am not summarizing: the full prompt, Claude's raw output, and how a human rejected it and asked again are reproduced as they happened.

impact is a small pipeline that takes a decision atom's ID and produces an impact graph. It merges the inbound edges (the decision's rationale), the outbound edges (files the decision will modify), the ontology's affects relations, and the wikilink reverse references (other documents pointing at this decision via [[D2026_Q2_017]]) into a list of impact candidates.

Prompt (Full Text)

Evaluate the proposal-stage pre-impact of decisions/D2026_Q2_017.md.
- Files this decision will modify (outbound)
- Documents that reverse-reference this decision as [[D2026_Q2_017]] (inbound wikilinks)
- Systems connected through the ontology's affects relations
Combine the three into impact candidates, and for each candidate, classify whether
it needs my confirmation (T1) or can be confirmed automatically. Don't fill gaps
with guesses; if there is no supporting file, mark it 'no evidence'.

Claude's Raw Output

I read decisions/D2026_Q2_017.md. Here is the impact graph extracted from the frontmatter.

Outbound (directly modified by this decision) - data/CombatFormula.md — the global_cooldown field (explicit in the atom) - sheets/CombatBalance — the healing-skill exception rows (explicit in the atom) - ui/ComboIndicator — inferred from the atom's affected_teams: [ui]. However, no specific file is in the atom → no evidence (inferred)

Inbound wikilinks (reverse references to this decision) - I searched with grep for [[D2026_Q2_017]], but the current index shows 0 reverse references. This appears to be a new decision that nobody has cited yet.

Ontology affects relations - The ontology contains the relation CombatFormula → affects → TankSurvival. - This decision therefore indirectly affects the TankSurvival metric. Estimated impact: +12%.

This is where the human stops. The last line, "estimated impact +12%," is the problem.

Human Verification and Rejection

Rejected. Where did "estimated impact +12%" come from? That number is in neither the atom nor the ontology. If you made it up, take it out. Report only the fact that the affects relation exists — the size of the impact comes from running a simulation. And if you classified the UI item as 'no evidence,' why did you put it in the candidate list as T1?

Claude's Follow-Up Response

You're right. I'm correcting two things.

  1. "+12%" withdrawn. There is no basis for that number anywhere in the atom or the ontology. It is a value I generated. I should report only that TankSurvival is affected, and the size should be filled in from the simulation (Layer 2) results. Its current value is unmeasured.
  2. UI item reclassified. Since I could not find a specific file, it should have been demoted from T1 (needs confirmation) to insufficient evidence — a human needs to specify the file. The corrected table is below.
Impact candidate Edge type Evidence Classification
CombatFormula.md global_cooldown outbound explicit in atom auto-confirmed
CombatBalance healing exception rows outbound explicit in atom auto-confirmed
TankSurvival metric affects ontology relation affected / size unmeasured
UI combo display (file unknown) inferred no evidence human must specify file
inbound reverse references wikilink grep: 0 hits not applicable

This table is the deliverable of pre-tracking. The key row is TankSurvival. The impact that broke PvP in the opening anecdote is already showing up in pre-tracking as "affected / size unmeasured." If we had seen this in the meeting, the response would have been "then let's measure the size with a simulation before we decide." Letting the AI make up +12% would have been more dangerous, not less. Fake precision makes people skip verification.


18.3.3 The Decision and the Irreversible Stages

Once pre-tracking is done, the meeting decides. The moment the decision is finalized as an atom, two irreversible stages begin.

flowchart TD
    A["Pre-tracking
(impact graph)"] --> B{Decision meeting} B -->|rejected| Z["Discard proposal
(reason logged, reversible)"] B -->|approved| C["Decision atom finalized
D2026_Q2_017"] C --> D["⚠ Irreversible 1: applied to build
edit CombatFormula & sheet"] C --> E["⚠ Irreversible 2: propagated to other docs
update GDD & reverse refs"] D --> F["Post-tracking begins"] E --> F F --> G{Prediction = measurement?} G -->|matches| H["Decision card closed"] G -->|deviates| I["Side-effect atom
follow-up decision candidate"] I --> B classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; classDef fail fill:#fee2e2,stroke:#dc2626,color:#7f1d1d; class D code; class B human; class A,C,E,I data; class H pass; class Z fail;

The reason they are irreversible is simple. A number applied to the build has already been played by users, and content propagated to other documents has already been used by teammates as the basis for their next tasks. That is why the portal_layer_change_impact_check atom sits as a gate right before these two stages. Its rule fits in one line: "If a decision includes a change that crosses Layers (data → formula → UI), a human must approve the impact candidate list once before it is applied." If only auto-confirmed items remain, it passes; if even one item reads 'human must specify file,' it blocks. The UI item from 18.3.2 becomes a blocking reason.


18.3.4 Post-Tracking — Attaching Measurements to the Same Graph

Post-tracking runs one week to one month after the change lands in the build. The job is to attach a measured value to each row of the impact candidate table that pre-tracking produced.

Prompt (Full Text)

Post-tracking for D2026_Q2_017. Attach a measurement to each candidate in the pre-tracking table (18.3.2).
Measurement sources:
- CombatBalance sheet git diff (to confirm the change landed)
- 1000-run simulation log sim_2026q2_017.json
- launch week-1 KPI export kpi_w1.csv
Compute the gap between intended and measured values, and collect separately only the rows outside the 10% error range.
Mark missing data as 'not collected'.

Claude's Raw Output

I combined the measurements with the pre-tracking table.

Candidate Intended Measured Gap Verdict
global_cooldown 0.3→0.5 applied confirmed via git diff applied
healing exception rows applied confirmed in sheet applied
combat duration +15% +14.2% -0.8%p within range
TankSurvival (unmeasured) +44% no baseline no verdict
5:5 PvP win rate +15%p +29%p +14%p deviation
UI combo readability improved 65% positive user feedback matches

Rows outside the error range (±10%): PvP win rate +29%p (intended +15%p). TankSurvival has no intended value because its size was not measured beforehand, so no verdict is possible.

Connecting the two rows: TankSurvival +44% appears to be the cause of the PvP win-rate deviation. Leaving TankSurvival as "size unmeasured" in pre-tracking foreshadowed this deviation.

Human Verification

Correct. Not measuring TankSurvival in the simulation is the root of the accident. Create a follow-up decision atom: a proposal to bring the PvP win rate back into the intended range with a dmg adjustment for the tank classes.

Claude's Follow-Up Response

I generated a draft of decisions/D2026_Q2_025.md. I set [[D2026_Q2_017]] on its inbound side, making it explicit via reverse reference that this decision is a side-effect correction of decision 17. Querying atom 17 with impact again now picks up one inbound reverse reference (number 25) — the slot that was 0 in 18.3.2.

This is where the graph closes. The node that was "affected / size unmeasured" in pre-tracking was confirmed as a deviation in post-tracking, and a follow-up decision came in as a reverse reference pointing at that node. The decision's entire cycle made one full loop on the same graph.


18.3.5 The Actual Commands Behind the Tracking — The grep Workflow

The inbound reverse-reference extraction in impact is not some fancy tool; it is one line of grep. It searches every document for wikilinks pointing at the decision atom.

# All documents that reverse-reference D2026_Q2_017 (inbound wikilinks)
grep -rln "\[\[D2026_Q2_017\]\]" decisions/ manuscript/ gdd/

# The decision atom's outbound edges — extract affected_files from the frontmatter
grep -A20 "affected_files:" decisions/D2026_Q2_017.md

# Post-tracking: only the rows that deviate from intent (the verdict column)
grep -E "이탈|판정 불가" tracking/D2026_Q2_017_post.md

Three lines run the skeleton of pre- and post-tracking. The LLM's place is to read and interpret these results, not to replace the search itself. grep supplies the facts (which files point at this decision), the LLM weaves those facts into an impact candidate table, and a human takes responsibility for the size of the impact and the verdict. This separation is why "don't make up +12%" worked in §18.3.2.


18.3.6 Measurement — When Pre- and Post-Tracking Are Tied Together

These figures compare my Project A before and after standardizing the decision cycle. The absolute time figures are the author's estimate (unverified), dependent on team size (mid-sized, 10–50 people); the ratios and directions were observed in actual operation.

Item Pre/post tracking separated Pre/post tracking unified
Share of decisions where post-tracking actually ran about 30% 90% or more
Impacts flagged pre that blew up as accidents post common almost none (gated up front)
Side effect → follow-up decision linkage rate low (passed along verbally) auto-candidates via reverse references
Completeness of inbound reverse references in the decision graph patchy closed loop

One thing matters most. When pre-tracking and post-tracking share the same candidate table, a hole left as "size unmeasured" up front gets checked in exactly that spot afterward. When they are separate, what was seen beforehand and what was measured afterward live in different formats, so they cannot be compared — which is why the tracking rate stalls at 30%. That said, targeting 100% reverse-reference completeness from day one only adds operational burden. The realistic path is to start with the habit of writing affected_files in decision atoms, and to fold the reverse-reference grep into your retrospective cycle, expanding gradually.


18.3.7 Common Failures

Pattern Remedy
Saw the impact up front but decided without measuring its size hold the decision on every "size unmeasured" row until the simulation runs
The LLM makes up impact numbers no supporting file → 'no evidence'; sizes come from simulation only
Post-tracking uses a different format from the pre table add only a measured column to the same candidate table
Side effects handed over verbally enforce a follow-up decision atom + reverse-reference wikilink
Layer-crossing changes applied without a gate make passing portal_layer_change_impact_check mandatory

Key Takeaways


Beyond Games. Reading twice — checking "how far will this spread?" before you commit a decision (pre) and "did it really spread only that far?" after you apply it (post) — is a basic move of all change management, not just games. When a company changes its pricing policy, putting the affected departments (sales, CS, billing) on a candidate table beforehand and leaving the size as "unmeasured until simulated" heads off the post-launch accident of "why didn't the billing team know about this?" For example, before introducing a new membership tier, add post-tracking columns such as CS ticket volume and churn rate to the pre table as blank cells; a month later you fill those cells with measurements and compare intent against reality directly in the same table.

Try It Yourself

setup — Create the decisions folder and the tracking folder.

mkdir decisions tracking
# In one decision atom, record affected_files and affected_teams in the frontmatter

prompt — Connect pre-tracking to post-tracking through the same table.

Pre-impact assessment for decisions/<ID>.md: combine the outbound edges (files to modify),
inbound wikilinks, and ontology affects relations into an impact candidate table,
marking items without evidence as 'no evidence' and sizes as 'unmeasured'. Don't make up numbers.

(after the change lands in the build)
Attach only a measured column to the same candidate table and collect the rows more than 10% off from intent.
Turn each deviating row into a follow-up decision atom draft and set a [[<ID>]] reverse reference.

verify — Confirm with grep that the graph has closed.

grep -rln "\[\[<ID>\]\]" decisions/   # loop closed if the follow-up decision's reverse reference shows up
grep -E "이탈|미측정" tracking/<ID>_post.md   # check for remaining gaps

Solo Scale-Down

If you are a solo game developer, drop the meetings, owners, and deadlines entirely. When you write a one-line decision in a decisions/ markdown file, fill in just two fields: affected_files: (the files this decision will touch) and expected: (the change you intend). After you build, open those files and check with your own eyes that things went as intended; if something is off, add one actual: line to the same file. One tool is enough: grep -rln "[[decision-ID]]". One field before, one field after — that is the minimal form of pre/post tracking.

18.4 The Document Impact Grep Workflow — Pulling the Impact Scope with impact

Monday, 10 a.m. Team member A, who owns combat, dropped one line into the team messenger: "Can we lower the global cooldown (GCD) from 0.5 seconds to 0.4?" It is a one-number change. On the surface. I read that line and my hands stopped. How many documents have this number entered in them, how many skill balance atoms were built on this constant as a premise, which sheets' formulas break if it changes — none of it came to mind. And believing it has come to mind is exactly how accidents happen. The misses that kept erupting at 8 to 12 per quarter — the "I never saw that document" kind — were precisely this illusion.

So I decided to stop memorizing the answer. Instead, I type one line.

impact combat_global_cooldown_constant

This chapter looks at what that one line spits out, raw and unedited. It shows that "extracting the impact scope" is not an abstraction — it is the concrete act of scraping three channels with grep: inbound edges, ontology affects, and wikilink backlinks.


18.4.1 The Impact Scope Arrives Through Three Channels

"What is affected if I change this atom?" is actually three questions. Mix the three and the answer blurs; separate them and each one falls out as a single line of grep.

First, inbound edges — who points at me. If atom A references atom B, that is an edge in the A→B direction. When you change B, the danger lies in the As pointing at B — the arrows coming into B. So you look at inbound, not outbound (who I point at). The shockwave of a change travels backward, up the arrows.

Second, ontology affects — what it semantically influences. This is the affects: field declared in an atom's frontmatter. Even when the name never appears in the text, it is a semantic link the designer declared in advance: "this affects that." It is the alias-and-synonym problem — the one grep cannot catch — entered by a human ahead of time.

Third, wikilink backlinks — documents that explicitly link to me in [[atom_id]] form. This channel is the most reliable, because it is not a coincidental word match but a link an author placed on purpose.

Here is how the three channels relate.

combat_global _cooldown_constant Inbound edges Who references me Ontology affects Declared affects: field Wikilink backlinks Explicit [[atom_id]] links Impact scope list Deduplicated · tier labels

impact is the tool that bundles the three channels into one. The next section shows the output this tool actually produces, in full.


18.4.2 Worked Transcript — Running impact Once

Instead of explaining in the abstract, I walk through one real cycle from start to finish: the full prompt, the tool's raw output (untouched), the parts I verified and rejected, and the re-request — all of it.

Prompt (Full Text)

I asked Claude to run the impact workflow. The workflow itself is a shell script that calls ripgrep, but interpreting the results and recommending tiers is Claude's job.

Run the impact workflow on combat_global_cooldown_constant.
Pull each of the three channels (inbound edges / affects / wikilink backlinks) with grep,
merge the duplicates, and attach an impact tier recommendation to each affected target.
The change: "global cooldown 0.5s → 0.4s".

Claude's Raw Output (Verbatim, No Summary)

I will run the three channels in order.

[1/3] Inbound edges — locations that reference this atom
$ rg -n "combat_global_cooldown_constant" --type md docs/ atoms/

atoms/combat/combat_skill_cooldown_rule.md:34:  base: combat_global_cooldown_constant
atoms/combat/combat_gcd_ui_feedback.md:12:  depends_on: combat_global_cooldown_constant
docs/spec/combat_v3.md:42:  based on the global cooldown (combat_global_cooldown_constant)
docs/spec/skill_design.md:18:  cooldowns are designed as multiples of combat_global_cooldown_constant
docs/balance/refgame_comparison.md:77:  refgame is 0.5s; our combat_global_cooldown_constant is the same

→ 5 direct references found.

[2/3] Ontology affects — places that name this atom in an affects field
$ rg -n "^\s*-?\s*combat_global_cooldown_constant" -B3 atoms/ | rg -B3 "affects:"

atoms/combat/combat_dps_curve.md:
  affects:
    - combat_global_cooldown_constant   ← (note: direction needs checking)

→ 1 hit. Note, however, that this is dps_curve declaring that it "affects" this atom,
  so the direction may be reversed. Human review needed — see the verification items below.

[3/3] Wikilink backlinks — [[...]]-style links
$ rg -n "\[\[combat_global_cooldown_constant" --type md

docs/spec/combat_v3.md:42:  ...[[combat_global_cooldown_constant]]...
atoms/combat/combat_skill_cooldown_rule.md:34: ...[[combat_global_cooldown_constant]]...

→ 2 hits. Both are locations already caught in [1/3] (duplicates).

[Merged — unique affected targets after dedup]
1. atoms/combat/combat_skill_cooldown_rule.md   (inbound+wikilink)
2. atoms/combat/combat_gcd_ui_feedback.md       (inbound)
3. docs/spec/combat_v3.md                        (inbound+wikilink)
4. docs/spec/skill_design.md                     (inbound)
5. docs/balance/refgame_comparison.md            (inbound)
6. atoms/combat/combat_dps_curve.md              (affects — direction unconfirmed)

[Tier recommendations — per the five impact tiers of 18.2]
1. combat_skill_cooldown_rule    → T1 (every skill cooldown is a multiple of this constant)
2. combat_gcd_ui_feedback        → T2 (UI feedback timing; visual recalibration needed)
3. combat_v3.md                  → T3 (spec doc number update)
4. skill_design.md               → T3 (design doc number update)
5. refgame_comparison.md         → T2 (benchmark comparison table; recompute the gap vs. refgame)
6. combat_dps_curve.md           → TBD (tier after direction check)

Human Verification / Rejection

Two places stopped me here.

Rejection 1 — the affects direction. combat_dps_curve.md listing our constant in its affects: field is a declaration that dps_curve affects this constant. But what we are changing now is the constant itself. The arrow points the other way. dps_curve does not change the constant; rather, when the constant changes, the premise of dps_curve's calculation is shaken. Claude was right to hold it honestly with "direction needs checking." I kept it as an affected target but reclassified it — not affects, but a "reverse-direction semantic dependency." Confirmed at T2.

Rejection 2 — rating the refgame comparison table T2 was an overestimate. refgame_comparison.md is a plain number-comparison table against an external game. If our number changes, you fix one cell from 0.5 to 0.4 and you are done; no system uses this table as input. Demoted to T3 (doc number update).

Re-Request

Two corrections — redo the impact table.
1) Keep combat_dps_curve.md as an affected target, relabel it "reverse-direction semantic dependency",
   tier confirmed T2.
2) Demote refgame_comparison.md from T2 to T3 (no system uses it as input; one table cell to fix).
Then pull out only T1 and T2 into a "must review before the change" list.

Claude's final result came back as follows.

[Must review before the change — T1·T2]
T1  combat_skill_cooldown_rule.md   every skill cooldown is a multiple of this constant
T2  combat_gcd_ui_feedback.md       recalibrate UI feedback timing
T2  combat_dps_curve.md             reverse-direction semantic dependency — recheck the DPS curve's premise

[Doc number updates only — T3]
T3  combat_v3.md / skill_design.md / refgame_comparison.md

The six affected targets that "did not come to mind" at the start became a prioritized list after one grep cycle and two human judgments. That is what impact-scope extraction actually is. The tool scrapes the candidates exhaustively; a human settles direction and tier.


18.4.3 The Extraction Pipeline — What Is Automated and What Is Human

Generalizing the previous section's cycle into a flow gives this. The key is where the automated steps and the human step part ways.

flowchart TD
    A[Specify the atom to change
combat_global_cooldown_constant] --> B{Run impact} B --> C1[Inbound edge grep
rg atom_id] B --> C2[affects field grep
rg affects: block] B --> C3[Wikilink backlink grep
rg bracketed atom_id] C1 --> D[Dedup and merge] C2 --> D C3 --> D D --> E[LLM tier recommendation
T0~T4 labels] E --> F{Human review} F -->|Direction error / tier inflated| G[Reject and re-request] G --> E F -->|Approve| H[Finalize pre-change review list
T1·T2 separated] H --> I[Auto-attach to the change request comment] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; class B,C1,C2,C3,D,I code; class E ai; class F,G human; class A,H data;

The automated parts are the three grep channels, the merge, and the tier draft. The human part is exactly one thing: the final call on direction and tier. The affects reversal and the refgame demotion from the earlier cycle happened precisely at this spot. Trust the tool 100% and you get two kinds of accidents: dropping the reversed affects from the affected list, or overprotecting a comparison table and running a needless review every single time. Placing the boundary between automation and human judgment at this one point is the design intent of the workflow.


18.4.4 The Three-Channel Grep Pattern Reference

Here are the ripgrep patterns impact calls internally, written out as is. This is what the tool really is — not fancy infrastructure, but three proven lines of regex.

Inbound edges. Every location where the atom ID appears in document text. The widest scrape.

rg -n "combat_global_cooldown_constant" --type md docs/ atoms/

The affects field. Only the cases where the atom ID sits inside an affects: block. -B3 pulls in the three preceding lines, so a human can eyeball whether it is an affects block or some other field.

rg -n "combat_global_cooldown_constant" -B3 atoms/ | rg -B3 "affects:"

Wikilink backlinks. Only explicit links wrapped in double brackets. The most reliable, so they go to the front of the review queue.

rg -n "\[\[combat_global_cooldown_constant" --type md atoms/ docs/

Across the three patterns, precision and recall trade off exactly. Wikilinks are nearly 100% precise but miss anything the author never linked. Inbound edges scrape everything but mix in coincidental word matches (noise). Affects captures meaning but muddles direction. Only together do they plug the holes. Use any one alone and something will always leak.


18.4.5 Tying into the Decision Card — portal_layer_change_impact_check

Impact-scope extraction is one stage of the decision cycle (§18.3). The moment a decision card is registered, impact runs with the card's affected_atoms slot as input. The atom that enforces this link is portal_layer_change_impact_check.

This atom's job is to make sure that when a change crosses Layers, the impact check cannot be skipped. The cooldown constant change is one number in L1 (systems), but its impact spreads to L3 (data sheet formulas) and L4 (build QA items). portal_layer_change_impact_check judges whether a change crosses a Layer boundary, and if it does, it forces an impact run.

---
name: portal_layer_change_impact_check
type: gate
description: A change that crosses a Layer boundary must not ship to the build before passing the impact check
trigger:
  - affected_atoms is non-empty when a decision card is registered
  - the changed atom's layer != the affected atom's layer
action:
  - run impact (three-channel grep)
  - if any T1·T2 affected targets exist, block the merge until the "review complete" box is checked
---

In the cooldown case, what this gate caught was not combat_skill_cooldown_rule (L1) but the CombatBalance sheet (L3) that takes the rule as input. The sheet's cooldown-multiplier column builds its formulas on the constant. Grep scraped the atom out of the documents, and the gate pushed back: "this crosses a Layer — check the sheet too." Without the two tied together, you get the classic miss: the documents get updated while the sheet keeps running on the old premise.


18.4.6 Measurement — What the Workflow Recovered

These are the changes observed in my Project A operations. The time figures are the author's estimates (unverified); the miss counts are values actually tallied in the quarterly retrospectives.

Item Without the workflow With impact in operation
Time to identify affected atoms relying on memory (incomplete) 1–2 minutes (exhaustive grep)
Missed-change accidents 8–12 per quarter (tallied, measured) 1–2 per quarter (tallied, measured)
Impact attached to change requests a human, occasionally enforced by the gate
New team member grasping impact days (oral handover) 30 minutes (tool + cards)
Infrastructure cost graph DB adoption under review ripgrep + shell only

The last row is the conclusion of this entire chapter. Project A evaluated a graph DB and a search index, and ended up settling on ripgrep and a small shell script. A precision instrument is, in fact, more accurate than a tape measure. But the tool you pull out every day converges on the tape measure — the one that does not break and needs no infrastructure. Misses dropped from 8–12 to 1–2 per quarter not because the tool is sophisticated, but because it runs every single time, without exception.


18.4.7 Limits — What Grep Cannot Catch

Even with the three channels bundled, leaks remain. Using a tool while knowing its limits is different from trusting it blindly.

Aliases and abbreviations. If a document writes only "GCD (Global Cooldown)," it never matches a combat_global_cooldown_constant grep. The mitigation is to expand the search term into a regex — (combat_global_cooldown_constant|GCD|전역\s?쿨다운). Keep a team abbreviation dictionary as a separate file and compose it into searches automatically.

The irreversible zone. Grep is a tool for the reversible stage. Before a change ships to the build, the impact among documents, atoms, and sheets is fully visible to grep. But what happens after the build goes out and players feel the 0.4-second cooldown — community complaints, a shift in perceived tempo — is not something grep can search. So the principle is simple: finish every grep review before the change ships to the build. Once you cross into the irreversible stage, what grep can tell you drops off sharply.

Where LLM review fits. Just as a human settled the affects direction and the tiers in the earlier cycle, inserting an LLM to judge the relevance of grep candidates filters out the noise. But the LLM is not 100% either, so final approval stays with a human. Accuracy reaches an operable level in a structure where tool, LLM, and human each filter one stage. Remove any one stage, and the kind of accident that stage was catching comes right back.


Beyond Games. The habit of finding "what shakes if I change this item" by exhaustive search instead of memory pays off the same way for any office worker who lives in documents and spreadsheets. When you revise one clause in a terms-of-service document, trying to recall from memory every contract, notification email, and customer FAQ that cites that clause number guarantees a miss; grep the whole folder by keyword, scrape it exhaustively, then have a human sort each hit into "must fix / flag only / unrelated," and the misses disappear. For example, when an accountant changes a particular account code, searching exhaustively for every settlement sheet, report template, and macro that references the code, and turning the hits into a pre-change review list, structurally prevents the quarterly closing accident of "I never saw that one sheet."

18.4.8 Try It Yourself

setup

If your documents and atoms live in plain text (.md) and ripgrep (rg) is installed, you are ready. An atom ID naming convention (snake case, unique IDs) raises grep precision considerably.

# Verify: how many times one atom ID appears across all docs
rg -c "combat_global_cooldown_constant" docs/ atoms/

prompt

Give it the atom you are changing and the change itself, and ask for the three-channel extraction plus tier recommendations.

Run impact on <atom_id>.
Pull each of the three channels (inbound edges / affects / wikilink backlinks) with grep,
merge the duplicates, then recommend an 18.2 impact tier (T0~T4) for each.
The change: <what, changed to what>.
Separate out only T1·T2 into a "review before the change" list.

verify

Do not take the tool's output on faith; check two things by hand.

  1. The affects direction — for each item caught via affects, check whether it is the side doing the affecting or the side being affected. If the direction is reversed, fix the label.
  2. Tier inflation/deflation — if a table or a comparison-only document comes up as T1·T2, ask: "Does any system use this document as input?" If not, demote it to T3.

After checking, attach only the T1·T2 list to the change request comment, and the cycle closes.

Solo Scale-Down

If you work alone — no tooling, no atom graph — one command and one memo column get you the same effect.

# Exhaustive search of the whole folder for the concept name you are about to change
rg -n "전역쿨다운|GCD|global_cooldown" .

Paste the search results straight into a notepad and, next to each line, write one of three labels by hand: "must fix / flag only / unrelated." That is the one-person version of impact. The point is not the tool's sophistication but the procedure itself: scrape exhaustively instead of relying on memory, then have a human classify. With the procedure, misses shrink; without it, that Monday-morning blankness repeats every time.


Key Takeaways

Next Chapter Preview

Part 19 · Team Lead

19.1 Turning the Vision into a Scorecard for Decisions — Running 26 decisions/ Atoms Through an LLM

Primary audience: design directors and lead designers running mid-sized (10–50 person) teams Scaled-down version for solo/hobbyist readers: §19.1.8 "If You're Solo, Just This Much"

Even teams that have written a solid one-page vision document repeat the same accident. The vision hangs on the wall, but nobody checks whether the decisions piling up every week actually match it. Someone pulls it out once at the quarterly retrospective, but by then three more decisions have already been stacked on top of the one that drifted. For a vision to become "the reference point in disputes," what matters is not the writing but running every decision against it. And that cross-checking is exactly the kind of work that is tedious and easy to skip when done by hand — a perfect job to hand to AI.

This chapter ties two things together. The front half is a workflow that turns an already-written vision into a scorecard for decisions — one full cycle of running the 26 actual decision atoms from my project through an LLM, receiving "vision slot violation" verdicts, and having a human catch one misjudgment among them. The back half is the question of whose decisions that scorecard covers — that is, delegation of authority. General leadership theory (why vision matters, why delegation is a growth tool) already fills other books, so this chapter stays in one spot: running those principles as an AI workflow.


19.1.1 Vision, Roadmap, Schedule — Just Enough on Why the Three Layers Differ

First, what does it mean for a vision to filter decisions? Vision, roadmap, and schedule are not the same thing. They differ in time unit and rate of change, and when that difference collapses, schedule pressure starts shaking the vision.

Layer Horizon Rate of change What checking against the vision means
Vision 5–10 years Almost never The baseline decisions must align with
Roadmap 1–3 years Quarterly The middle layer that translates vision into a timeline
Schedule 1–3 months Weekly Not checked against the vision directly

The key point is that decisions are checked against the vision — the layer that changes least. When the schedule is tight, you do not change the vision; when the schedule conflicts with the vision, you fix the schedule. This hierarchy has to be firm for the automated check in the next section to mean anything. If the baseline of the check wobbles every week, the check itself is pointless.

The vision is one page, five slots, done. Here is the skeleton of my project's vision document. These slots become the scoring criteria for the LLM check in §19.1.3, so take in the shape first.

---
title: Project A Vision v2
layer: L0
locked: true   # Changes require game director + CEO agreement
---

## Slot 1. What We Are Building
A mobile-first MMORPG set in a Korean fantasy world.

## Slot 2. Who It Is For
Users in their 30s to 50s, mostly on mobile, who enjoy serious storytelling.

## Slot 3. Why (Differentiation)
- Deep storytelling through a multi-layered narrative (depth, not mass production)
- Simultaneous operation in Southeast Asia + Korea

## Slot 4. How (Values)
- Respect users' time (minimize filler content)
- Decisions balance data + people
- Team consensus takes priority over decision speed

## Slot 5. What We Are Not
- Not a runaway F2P monetization model
- Not PvP-centric
- No forced N hours of daily attendance

Slot 5 ("What We Are Not") does the most work in the check. Violations usually come not from "what we decided to do" but from quietly doing "what we decided not to do."


19.1.2 The Decisions Are Already Stacked Up as Atoms

What do you run the vision against? My team pins every major decision down as one atom each in the decisions/ folder. Each is a factual record with date, parties, and rationale spelled out, and 26 of them have accumulated so far. The input to the check is these 26 — you are not creating anything new, you are running what already exists.

Here is the actual shape of one decision atom (anonymized).

---
type: decision
id: D0019
date: 2026-05-12
deciders: [game director, data director]
tier: T1
---
# refgame_selective_adoption_for_mobile
Selectively adopt part of the reference MMORPG's combat data for the
mobile build.
Rationale: the combat pacing is proven on a 6-inch mobile screen, and
redesigning from zero would push the alpha schedule back by a quarter.
However, the monetization and attendance-incentive structures are not
adopted.

From the 26, here are a few representative ones used as check input (actual atom names, §A.3.3).

atom id atom name (anonymized) tier One-line summary
D0007 claude_role_transition_phase2 T1 Promote Claude from passive assistant → active partner
D0014 dataset_scope_alpha_split T2 Finalize the alpha dataset split criteria
D0019 refgame_selective_adoption_for_mobile T1 Selectively adopt reference game combat data
D0021 procedural_capability_frontier_5stage T1 Define the 5 stages of procedural generation capability
D0023 class_keyword_world_only T2 Restrict class keywords to the game world

This table is the input data for the prompt in the next section. Running all 26 at once is the point: if a human cross-checks 26 decisions against the vision one by one at a retrospective, it takes half a day, focus fades from the midpoint, and violations slip through. That tedious first pass is what gets handed to the LLM.


19.1.3 [Worked Transcript] Running 26 Decisions Against the Vision

Let's watch one cycle all the way through. The input prompt can be copied as is; the output is a reconstruction of an actual session.

Step 1 — The Prompt: Hand Over the Vision as a Scorecard and Force a Verdict Format

Run each of the 26 decision atoms below against the 5 vision slots, and
judge whether each decision is aligned, complementary, a violation, or
unrelated. For violations and unrelated ones, add a one-line rationale
saying which slot and why (pay special attention to Slot 5, "What We
Are Not"). Mark anything ambiguous as "ambiguous" and pass it to me.
No rationale needed for aligned ones. Give me a table only
(id | verdict | conflicting slot | rationale | ambiguous), no overall
commentary.

[Vision 5 slots — paste the full text from §19.1.1]
[26 decision atoms — id, body, rationale]

Half of this prompt is "forcing the verdict format." Given free rein, the LLM smears all 26 into "broadly aligned." The four-label verdict scheme, the rationale requirement only for violation and unrelated, and the escape hatch "if it's ambiguous, pass it to a human" are the heart of this prompt. Naming Slot 5 explicitly in the rationale instruction is deliberate too — violations come from "what we decided not to do."

Step 2 — Claude's Output (Reconstructed, Excerpt)

| id    | Verdict   | Conflicting slot | Rationale                                  | Ambiguous |
|-------|-----------|------------------|--------------------------------------------|------|
| D0007 | Unrelated | —                | Claude ops decision. Different axis from the game vision |      |
| D0014 | Unrelated | —                | Dataset split. Internal infrastructure decision |      |
| D0019 | Violation | Slot 5           | Adopts reference game combat data → possible | ★    |
|       |           |                  | conflict with "not a runaway F2P monetization model" |      |
| D0021 | Aligned   | Slot 3           |                                            |      |
| D0023 | Aligned   | Slot 1           |                                            |      |
| ...   |           |                  |                                            |      |

(Of 26: 18 aligned · 3 complementary · 3 unrelated · 2 violations)
Violation/ambiguous review requested: D0019, D0026 — human review needed

The most valuable part of the output is not the table but the bottom: the two violations and the ambiguity flags. The LLM automatically filtered out 24 of the 26 and raised only the 2 that a human needs to look at. A half-day cross-check shrank to reviewing 2 items. But one of those 2 is a misjudgment.

Step 3 — Verification and Veto (the Human's Seat)

A human rereads the D0019 (refgame_selective_adoption_for_mobile) verdict. The LLM saw "adopts reference game combat data" and judged it to conflict with Slot 5's "not a runaway F2P monetization model." On the surface, the words look plausible — the reference game is famous for aggressive monetization.

But read the atom body to the end and there is a final sentence: "However, the monetization and attendance-incentive structures are not adopted." The decision takes only the combat pacing data and explicitly excludes the monetization structure. If anything, it is a decision that upholds Slot 5. The LLM failed to give that final qualifying sentence its weight in the verdict and, pulled by the source word "reference game," classified the decision as a violation. This is not a Slot 5 violation — it is aligned.

Why this kind of misjudgment happens is clear. The LLM weighs the decision's source (which game it came from) and the decision's content (what was taken and what was discarded) equally. A human knows that the qualifying clause "however, we do not adopt X" is the heart of the decision. So the human vetoes and re-requests.

Look at D0019 again. The last sentence of the body — "However, the
monetization and attendance-incentive structures are not adopted" — is
the key. Split what is adopted (combat pacing data) from what is
excluded (monetization/attendance structures), and judge again which
slot each side falls under.

The LLM answered again: "The adopted part (combat data) aligns with Slots 1 and 2; the excluded part (monetization structure) actively supports Slot 5. Overall verdict: aligned. The previous violation verdict was an error — an overreaction to the source word." With this one round trip, D0019 was corrected from violation to aligned. The remaining genuine review item is one: D0026.

This cycle is the heart of this chapter. The LLM reduces 26 items to 2, but one of those 2 can be a misjudgment. The automated check does not eliminate human review; it is a tool that lets a human focus on 2 items instead of 26. If a human does not read those 2 to the end, a perfectly sound decision goes to a meeting labeled "vision violation" and sparks a pointless dispute.


19.1.4 The Check Flow at a Glance

Keep the cycle above as a diagram, and the same flow repeats every quarter from then on. The key point is that an LLM verdict never overturns a decision automatically. Only violations and ambiguous items go up to the human gate, and discarding, correcting, or approving is done by a human.

flowchart TB
    A["Vision 5 slots (L0, locked)
Scoring baseline"] B["26 decisions/ atoms
id, body, rationale"] A --> C["LLM first-pass verdict
aligned/complementary/violation/unrelated + rationale"] B --> C C -->|24 aligned, complementary, unrelated| F["Pass — logged in quarterly retrospective"] C -->|2 violation, ambiguous| D{"Human review gate"} D -->|Misjudgment confirmed| E["Re-request
(reinforce qualifier and context)"] E --> C D -->|Genuine violation| G["Convene a decision re-review meeting"] D -->|Aligned after correction| F classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; class C ai; class D,E,G human; class A,B data; class F pass;

Human hands touch only two places: where vision and decisions are entered cleanly (the top), and the gate where the few items the LLM raised as violation or ambiguous are read to the end and judged (the middle). The tedious 26-item cross-check in between is run by the LLM. It is the same design as the city generator in §6.2, where lint did not auto-discard violations but only raised alerts to the writer gate — the machine picks the suspects, and a human decides whether they live or die.


19.1.5 Whose Decisions Does the Vision Check Cover — Delegation of Authority

A natural question follows: did the game director make all 26 decisions? They had better not have. If the lead makes every decision personally, they become the bottleneck; if they delegate everything, the vision weakens. The vision check is also a safety net for running delegated decisions through the same scorecard.

Decisions have tiers, and the tier is the authority. Here is my team's authority matrix.

Tier Decider Reviewer Notified Vision check target?
T0 vision/core Game director + CEO All leads Whole team The vision itself (the check baseline)
T1 system/cross-discipline TF chair + game director TF members Discipline teams ✅ Required
T2 discipline/mid-level Discipline director Seniors Discipline team ✅ Required
T3 one-off/small Senior Owner Directly affected Spot check
T4 immediate/hotfix Owner Senior (after the fact) Game director (after the fact) Excluded from the check

Most of the 26 in decisions/ are T1 and T2 — delegated decisions. The game director does not personally look at every T2. Instead, the vision check (§19.1.3) runs the delegated T1/T2 decisions against the vision once a quarter. The safety net of delegation is the vision check itself. T0 is not a check target but the check's baseline, and T4 hotfixes are excluded because they are high-volume and have almost no vision impact.

Delegation itself does not go to full autonomy in one move; it advances gradually through four stages.

Stage Authority Relation to the LLM check
1. Direct instruction "Do it this way" The delegator decides; no check needed
2. Advice + decision report "Decide with X in mind" Vision cross-check reviewed together at reporting
3. After-the-fact report "Decide, then tell me the outcome" Atom pinned → included in the quarterly check
4. Autonomous decision No reporting duty (within tier limits) As long as an atom is left, the check covers it after the fact

Stage 4 autonomous decisions carry the greatest risk of drifting from the vision — and that is exactly the risk the §19.1.3 check catches after the fact. Even a T2 decision made autonomously gets caught by the quarterly check automatically, as long as it is pinned as an atom. This is why the freedom of delegation and the consistency of the vision do not collide — decide freely, but the decision is left as an atom, and the atom is run against the vision every quarter.


19.1.6 Handling Numbers Honestly

A vision-and-delegation chapter faces a strong temptation to insert tables like "after adopting the vision, meetings dropped from 90 to 45 minutes" or "after delegating, the director's decision load fell from 30 a week to 5." Numbers like that, unverified, erode the book's credibility. The numbers in this chapter are handled in exactly one of three ways.

First, countable things are recorded as measured. The decisions/ atoms currently number 26 (measured as of May 2026). The number of items the LLM's first pass raised to the human gate, and the number corrected as misjudgments, are measured values counted from session logs. That 1 of the 2 violation verdicts in the worked transcript above (D0019) was a misjudgment is also the result of an actual session.

Second, effects are stated as direction only. "A half-day cross-check shrank to reviewing a handful of items" is the direction of the structure, not an absolute time. The exact time saved varies with decision count, team size, and atom body length, so the right reading is the structural difference between "26 by hand" and "LLM first pass + human gate." Outcome metrics like meeting time or morale scores are not driven by the vision alone, so I do not assert causation.

Third, only the measurable is promised. What this workflow can actually measure: the number of decisions run through the vision check per quarter, the number raised to the human gate, the misjudgment rate (the share of LLM violation verdicts a human corrected to aligned), and the number of unpinned decisions (delegated decisions that escaped the check because no atom existed). These four can be spoken in meetings as numbers, not "feelings." The misjudgment rate in particular proves, with a number every quarter, why LLM verdicts must not be taken at face value.


19.1.7 Common Failures

Pattern Why it fails Remedy
Writing the vision but never running decisions against it The vision stays wall decoration and decisions go their own way Run the 26 atoms against the vision every quarter (§19.1.3)
Taking LLM violation verdicts straight to a meeting Misjudgments (like D0019) spark pointless disputes For violation/ambiguous items, a human reads the atom body to the end
Not leaving decisions as atoms Delegated decisions drop out of the check entirely Make atom pinning mandatory in after-the-fact reporting (delegation stage 3)
Checking even T4 hotfixes Volume grows while vision impact stays near zero Limit check targets to T1/T2
Skipping from delegation stage 1 to 4 Autonomous decisions accumulate out of line with the vision Gradual delegation + quarterly checks as after-the-fact cover

The third is the one most often missed. The more smoothly autonomous a team is, the more it agrees on decisions verbally and never pins the atom. Then the §19.1.3 check sees only pinned decisions, and the most freely made decisions fall into the check's blind spot. The freedom of delegation is safe only on the premise of atom pinning.


Beyond Games. Running the vision against every decision, and delegating authority, are not homework unique to game teams — they are every manager's job. Pin your department's mission down as one page with five slots ("what we do / for whom / why / how / what we are not"), and once a quarter you can have an LLM do a first-pass cross-check of whether the practical decisions piling up each week have drifted from that mission — it is especially good at catching violations where something you decided not to do creeps back in quietly. For example, running the decisions a team lead has delegated through the department mission each quarter becomes a safety net that catches, after the fact, autonomous decisions that strayed off course. Just do not take the items the LLM flags as "violations" straight into a meeting — one of them may be a misjudgment, so a human has to read them to the end.

19.1.8 Try It Yourself — One Step You Can Take Today

If You're Solo, Just This Much: You do not need a decision atom folder. Write the vision for your own project (or hobby game) as a single page using the five slots from §19.1.1. Then jot down the last 5–10 decisions you made as one-line memos, paste the prompt from §19.1.3 as is, and run them through an LLM once. If even one "violation" verdict comes back, reread that decision's memo to the end and argue back yourself about whether the LLM is right. That is how you internalize what bundle of judgments a vision check really is — and why LLM verdicts must not be taken at face value.

If you are on a team, start with this one step. Lock the vision as one page with five slots (L0, locked), and pin each T1/T2 decision from the latest quarter as one atom in a decisions/ folder. With even 10 atoms accumulated you can run the §19.1.3 prompt once, and if that first cycle catches a single delegated decision out of line with the vision, the value of this workflow shows immediately.


Key Takeaways

Next Chapter Preview

19.2 Classify Conflicts and Don't Let Meeting Decisions Slip Away — AI Assistance for Meeting Leadership

Primary audience: directors and team leads who make 50+ decisions per quarter in meetings (midsize teams of 10–50) Scaled-down version for solo/hobbyist readers: §19.2.8 "If You're Solo, Just This Much"

I once ran a meeting well for 90 minutes, only to watch the same agenda item land back on the meeting table a week later. We had clearly decided — but who owned what was written down nowhere. The minutes said only "Discussed global cooldown," and "settled on 0.5 seconds, owner: Team Member A" evaporated from the heads of the people in the room within a week. The place where a leader's meetings collapse is mostly not during the meeting, but in the short gap right after it ends, before the decision hardens into a record.

This chapter covers two chunks of a team lead's job. The first half is how to route conflicts to standard remedies by type instead of solving each one from zero; the second half is the spine of this chapter — a worked transcript in which AI extracts the decisions from a meeting but is forced to block any whose owner or rationale is empty. General leadership theory (casting a vision, listening, empathy) is covered well enough in other books, so this chapter sticks to one spot: where that theory turns into an AI workflow that prevents decision leakage.


19.2.1 The Goal Isn't Zero Conflict — It's Classification

It's a misconception that a zero-conflict team is a healthy team. If a midsize team of 10–50 people makes more than 50 decisions a quarter and friction never once shows, the conflict isn't absent — it has sunk below the surface, and sunken conflict is more dangerous.

A leader's job is not to eliminate conflict but to classify it quickly by type and route it to a standard remedy. If the same conflict gets resolved a different way every time, the time it takes to resolve starts piling up from zero every time.

Conflict Type What's Actually Colliding Standard Remedy
Value conflict Differing readings of the vision (revenue vs. user time) Cite the vision slots
Fact conflict Different interpretations of the same data Check the data (metagame report)
Priority conflict "My area matters more" Compare impact grade and KPI impact
Authority conflict "This is my call" Recheck the authority matrix
Personal conflict Relationships and communication styles 1:1s, separate facts from feelings (outside the system)

For the first four types, the remedy is citing the system. When the vision, data, KPIs, and authority matrix are written down, the weight of a decision shifts from people's mouths to the system, and the debate gets shorter. Only the fifth, personal conflict, sits outside the system — 1:1s and separating facts from feelings; almost no tool besides time and sincerity works there. But "the system can't solve it" is no license for the leader to let go. The territory the system can't solve is still the leader's job — that is what makes this seat hard.

Classification doesn't restart from scratch each time; it runs as one loop.

flowchart LR
    A["Recognize conflict"] --> B["Classify type
(5 types)"] B --> C["Apply standard remedy
vision, data, KPI, authority, 1:1"] C --> D["Follow-up check after 1 week"] D --> E{"Recurs?"} E -->|"Same type repeats"| F["Review system and rules
(quarterly retro conflict slot)"] E -->|"Resolved"| G["Closed"] F --> A classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; class A,B,C,D,E,F human; class G pass;

The key is the branch on the right. When the same type of conflict keeps recurring, that's not a people problem — it's a system problem. At that point, instead of mediating between people, you fix the rules: the vision, the authority matrix. This becomes the input to the quarterly retro conflict slot covered in §19.2.7.


19.2.2 Meetings Exist to Make Decisions, and Decisions Must Not Slip Away

Just as four of the five conflict remedies are all "cite the system," a meeting, too, is ultimately a device that produces decisions and hardens them into records. The five principles a leader must hold in meetings are interlocked. Drop any one and the rest wobble with it.

  1. Share the agenda 24 hours before the meeting. (If people gather unprepared, the meeting drifts into an open-ended discussion)
  2. Enforce a time limit on each item. (5 minutes for information sharing, 15–20 for decisions, 30–45 for discussion; anything over carries forward)
  3. State "today's decisions" explicitly at the end of the meeting. (End without decisions, and the next meeting reopens the same items)
  4. The minutes are generated the moment the meeting ends. (Let a person "write them up later" and they evaporate)
  5. Track owner, rationale, and follow-up actions for every decision. (An untracked action disappears before the next week)

The incident from the opening was principles 3, 4, and 5 collapsing. The decision was made out loud (principle 3, partially met), never hardened into a record (principle 4, failed), and no owner was entered (principle 5, failed). So the same item came back a week later.

The problem is that if you leave principles 3, 4, and 5 to human willpower, they are the first to collapse in a busy week. The moment a meeting ends, the leader is already running to the next one. So we move these three principles into an AI-assisted pipeline. Decisions are extracted automatically from the minutes text, but anything with an empty owner or rationale is not allowed to pass. This pipeline takes the meeting → minutes → atom extraction flow built in Part 17 (§17.2) and looks at it once more from the leader's seat.


19.2.3 [Worked Transcript] Extract Decisions from the Minutes, but Block Any Without an Owner

Here is one full cycle of how this actually runs. The scene is right after a combat TF meeting on my project (a mobile-first MMORPG, "Project A" hereafter). The input prompts can be copied as is; the outputs are reconstructed from an actual session.

Step 1 — Input: Throw In the Raw Minutes as They Are

Don't pretty up the minutes. The input is rough text — utterances interleaved, lines that may or may not be decisions left as they are. Cleanup is the AI's job, not something a person does first.

[2026-06-05 Combat TF meeting minutes — excerpt, unpolished]

Team Member A: The sim results for taking the global cooldown to 0.5s came out stable.
Team Member B: If we lock healing skills to 0.5s too, I think the healing cycle breaks.
Team Member A: Let's carve those out. Healing gets a GCD exception.
Minsoo Lee: Sounds good — unify the GCD at 0.5s, healing is the exception. A, please
        do a full pass over the cooldown column in the data sheet.
Team Member C: For the targeting priority rules, let's look further next week and decide then...
Team Member B: The minimap zoom toggle probably needs the UI team in the room. On hold for now.
Minsoo Lee: Right, that goes to the next meeting.

Mixed in here are two decisions (GCD 0.5s, healing exception) and two holds (targeting, minimap). When a person picks them out by eye, one slips through each time. That was the incident from the opening.

Step 2 — Prompt: Demand Extraction, but Forbid Blank Owner and Rationale Fields

From the attached meeting minutes, pull out only the "decisions". Discussion,
holds, and information sharing are not decisions.
For each decision, fill four fields — decision (one sentence) / owner / rationale / follow_up —
but if you can't find the owner or rationale in the text, do not guess;
write "[MISSING — not finalized in the meeting]" instead. Put holds and items carried
over to the next meeting under deferred, and lines that may or may not be decisions
under ambiguous, and hand those to me. Output only the three blocks:
decisions / deferred / ambiguous.

Notice that half of this prompt is "blank enforcement." Give the AI freedom and it will invent a plausible owner or promote a hold into a decision. The escape hatch — "if you can't find it, don't guess; report [MISSING]" — is the heart of this workflow. A decision has value only when a person declares it explicitly (the principle of §17.6.3); the AI goes only as far as exposing that something is empty.

Step 3 — Claude's Output (Reconstructed from an Actual Session)

decisions:
  - decision: Unify the combat global cooldown at 0.5 seconds.
    owner: Team Member A ("A, please do a full pass over the cooldown column in the data sheet" — Minsoo Lee)
    rationale: "The sim results for taking the global cooldown to 0.5s came out stable" (Team Member A)
    follow_up: Full pass over the cooldown column

  - decision: Exclude healing skills from the global cooldown.
    owner: [MISSING — not finalized in the meeting]
    rationale: "If we lock healing to 0.5s too, I think the healing cycle breaks" (Team Member B)
    follow_up: Reflect the healing-skill cooldown exception in the data

deferred:
  - Targeting priority rules (Team Member C, further review next week)
  - Minimap zoom toggle (needs UI team consultation, on hold)

ambiguous: (none)

The most valuable part of this output is not the clean table of the first decision but the second decision's owner: [MISSING]. The healing exception was clearly decided, yet nowhere in the minutes did it say who would reflect it in the data. The AI didn't fill that hole with a guess; it reported it honestly. A good prompt lets the AI say "this slot is empty."

Step 4 — Verification and Rejection (the Leader's Seat)

Don't accept this output as is. A [MISSING] flag means the meeting finished only half of a decision. The leader's job here is not to fix the AI's output but to make the decision the meeting left unmade.

At this point I asked Team Member A one line over the team messenger: "You're also covering the data work for the healing exception, right?" A answered "Yes." That one line confirms the missing owner. Then I re-request.

The owner for the second decision (healing exception) is confirmed as Team Member A
(confirmed with them directly over the team messenger). Reflect this and give me
decisions again, and also convert both decisions into pending atom candidate format.
// (intent: include status: pending, source_meeting, owner, related_atoms — §17.2.4 format)

The AI answered again with the two owner-filled decisions converted into two pending atom candidates. These candidates do not become official decisions right away; they go through a one-week verification period in pending status (§17.2.4), because what a meeting settles sometimes gets overturned after a week of operation. It gives the ink time to dry. The cycle — input → extraction → MISSING flag → human completes the decision → re-request — closes here.

This one loop structurally prevents the incident from the opening. When a decision is only half made, that fact surfaces right there, right after the meeting — not a week later.


19.2.4 The Full Pipeline — Human Hands Touch Only Two Places

Lay the worked transcript above on top of Part 17's minutes pipeline and the whole picture looks like this. The leader's hands touch only two places: the seat where decisions are declared in the meeting (the very front), and the seat where the [MISSING] flags the AI raises get filled (the middle). The extraction, conversion, and registration in between are automatic.

flowchart TB
    A["Run the meeting
(leader: declare decisions out loud)"] --> B["Minutes text
(raw, unpolished)"] B --> C["AI decision extraction
4 required fields + MISSING flagging"] C --> D{"owner/rationale
empty?"} D -->|"MISSING"| E["Leader fills the gap
(messenger check → owner confirmed)"] E --> C D -->|"All filled"| F["pending atom candidates
(1-week verification period §17.2.4)"] F --> G["Weekly review
promote, discard, or hold §17.2.5"] G --> H["Register in JIT manifest
auto-injected next session §17.2.6"] classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; class C,D ai; class A,E,G human; class B,F,H data;

What the AI does not do in this pipeline matters more. The AI does not make decisions. It does not invent owners. It does not promote holds into decisions. What it does is pick decision candidates out of the minutes and expose the blanks — and there it stops. Declaring the decision and filling the blanks are done by people. This is the leader's-eye application of the principle from §17.6.3, "no AI auto-generation for decision slots" — once a decision propagates into other documents, sessions, and builds, it leaves irreversible traces, so at the entry gate we preserve a seat where a person declares it explicitly.


19.2.5 Enforcing [MISSING] Underpins the Equal Decision Culture

Among the team-shared atoms on the office PC is a concept atom named team_equal_decision_culture. It pins down vocabulary that kept recurring in retrospectives, and it names in one phrase a team culture: "decisions are made by evidence, not by job title." Not the director pressing "I decided, end of story," but every decision keeping its who and why, so that anyone can later retrace the decision from its evidence.

The [MISSING] enforcement of §19.2.3 is exactly the technical backbone of that culture. Refusing to let owner and rationale pass as blanks means a decision's authority rests not on "the director said so" but on "it came from this utterance in the minutes." Because a decision is blocked when its evidence citation is empty, a decision pushed through by title structurally cannot become an atom.

This culture also connects in a straight line to the conflict remedies of §19.2.1. Resolving value conflicts by citing the vision, fact conflicts with data, authority conflicts with the matrix — these are all the same principle: resolve with recorded evidence instead of someone's mouth. The equal decision culture is the soil the conflict remedies grow in, and [MISSING] enforcement is the tool that works that soil at every single meeting so it never hardens.

On top of this sits another axis of team culture: the boundary between open and closed. Meeting minutes, decision cards, KPI data, and incident reports live in the open zone; 1:1 conversations, performance reviews, salary, and personal circumstances live in the closed zone. Everything the decision-extraction pipeline handles is in the open zone. The same reason explains why personal conflict (the fifth type in §19.2.1) sits outside the system — it belongs to the closed zone, so it is never pinned into an atom.


19.2.6 How to Handle Numbers Honestly

A leadership chapter comes with a strong temptation to drop in a table like "we adopted the meeting pipeline and meeting time fell by half." Numbers like that, left unverified, cut into the book's credibility. This book's principle is one of three.

First, promise only measurable things as metrics. What the meeting pipeline can actually count is this: the number of missing owner/rationale fields per decision (target: 0), the share of decisions extracted from minutes that get promoted to pending atoms, and the number of "didn't we already decide this?" repeat meetings. These three can be spoken in a meeting as numbers, not feelings.

Second, label an estimate as an estimate. The claim that decision extraction right after a meeting takes "manual minutes cleanup, 20–30 minutes → AI draft plus gap-filling, under 5 minutes" is an estimate based on my experience — an unverified hypothesis. Don't memorize the absolute values; read the structural difference ("a person picks everything out from scratch" vs. "AI extracts, a person fills only the blanks"). The exact time saved varies with meeting size and decision count.

Third, don't assert causation. I don't nail down that "repeat meetings went down" is entirely thanks to this pipeline. Team maturity and project stage are at work too. State only the direction (when a decision omission surfaces right after the meeting, the system works toward fewer repeat meetings), and don't invent a multiplier.


19.2.7 The Conflict and Decision Slots in the Quarterly Retrospective

The conflict remedies and the decision pipeline run one inspection cycle in the quarterly retrospective. The retro carries a "conflict slot" and a "missed decision slot."

2026 Q2 quarterly retrospective — conflict & decision slots
─────────────────────────────────
[Conflicts] Top 3 this quarter
1. Global cooldown (value conflict) → closed by citing the vision.
   Learning: reconfirmed that the five vision slots work as decision criteria.
2. New dungeon priority (priority conflict) → compared KPI impact.
   Learning: no priority table, so we compared ad hoc every time → introduce a table next quarter.
3. Character design authority (authority conflict) → rechecked the authority matrix.
   Learning: the matrix needs a 'visual vs. functional' division-of-labor item.

[Missed decisions] [MISSING] occurrences this quarter
- Healing exception decision, owner not recorded (2026-06-05) → fixed over the team messenger.
  Learning: add "name the owner on the spot when a decision is declared" to the TF meeting facilitation checklist.

Conflicts and missed decisions are both inputs to the retro. When the same type of conflict recurs, you fix the system (the vision, the authority table); when [MISSING] keeps appearing in the same pattern, you fix how the meeting is run. The "review system and rules" branch that exited to the right in the §19.2.1 flowchart takes concrete form here.


Beyond Games. The meeting failure of "we clearly decided, yet the same item is back a week later" doesn't care what industry you're in. Feed your raw meeting notes to an LLM without polishing them, have it extract only the decisions, and have it report [MISSING] instead of guessing when the owner or rationale is empty — and the fact that a decision was only half made surfaces on the spot, right after the meeting. For example, in a weekly sales meeting, if "A will take this account" is only spoken and never recorded, it floats loose the following week; but if the AI extraction raises owner: [MISSING], you confirm the owner with a one-line message right then and erase one repeat meeting. The division of labor is the core: people declare decisions and fill the blanks, AI does the extraction.

19.2.8 Try It Yourself — One Step You Can Take Today

If You're Solo, Just This Much: You don't need a team or a minutes pipeline. Take the notes from a recent meeting you attended (a study group, a club, even a one-person project discussion all count) and paste them into the prompt from §19.2.3 as they are, then run it once. If even one decision comes back with owner: [MISSING], that's the item your team (or you yourself) will be reopening a week from now. Filling that blank now is enough to make one repeat meeting disappear.

If you have a team, start with this one step. Put your next meeting's minutes — unpolished — into the extraction prompt from §19.2.3, and keep only rule 2 (the [MISSING] enforcement) alive. Pending atoms and JIT registration (§17.2) come after that. Even that single blank-enforcing line catches the most expensive omission — the decision you thought you made but never wrote down — right after the meeting.


19.2.9 Common Failures

Pattern Why It Fails Remedy
Solving every conflict the same way No type ever gets resolved all the way Classify into 5 types → per-type remedy (§19.2.1)
Being content with a zero-conflict team Conflict sinks below the surface (more dangerous) Conflict is a health signal; quarterly retro slot
Making decisions out loud only, never writing them Same item re-meets a week later AI extraction + pinning as pending (§19.2.3)
AI fills in owners by guessing A wrong owner hardens into an atom Enforce [MISSING], forbid guessing (§19.2.2)
AI auto-generates decisions The decision's authority drifts away from its evidence People declare decisions; AI only reinforces (§17.6.3)
Promoting a hold into a decision An unconfirmed item propagates irreversibly Separate it into the deferred block (§19.2.3)

The third and fourth most often blow up together. A team that doesn't write decisions down hands the whole job to the AI — "just tidy this up for me" — and the AI helpfully invents an owner. Once that invented owner hardens into an atom, a week later you get a more expensive conflict: "I never agreed to take that." The [MISSING] enforcement blocks both failures with a single line.


Key Takeaways

Next Chapter Preview

19.3 AI Adoption Strategy and Executive Buy-In — From Conservative to Progressive, and No Doctored ROI

Primary readers: leads who have to decide whether to adopt AI on their team and explain the cost to executives (mid-sized teams of 10–50) Scaled-down version for solo/hobbyist readers: §19.3.12, "If You're Solo, Just This Much"

I once got asked, in the CEO's office, "We're paying this much a month for AI tools — so what exactly got better?" What I had in my hand was a single slide, and it said "Productivity up 3–5x." The CEO asked again: "Where does that 3–5x come from?" I had no answer. The number was a blog average I had copied from somewhere, not a value measured on my team.

Since that day I have stripped every doctored figure out of AI adoption reports. Instead I started reporting exactly what the system actually leaves behind — how many atoms have accumulated, how many skills are running, which inputs pull which context according to the logs. This chapter covers two things. First, a frame for deciding AI adoption in stages, from conservative (humans decide, AI verifies) to progressive (AI generates candidates, humans adopt). Second, how to explain the ROI of that adoption to executives with measured logs from my own system, not blog averages. General leadership theory is well covered in other books, so this chapter stays narrowly on using AI to assist the adoption decision itself, and drawing the evidence up out of system logs.


19.3.1 Adoption Is a Series of Stages, Not an On/Off Switch

If you treat AI adoption as a binary — adopt or don't — the first button goes into the wrong hole. Turn on five tools at once and the operational burden arrives before the benefits; too scared to turn anything on, and you never start at all. Adoption is a staged decision: start where the risk is low, and widen the authority as verification accumulates.

The criterion that runs through this entire book applies here unchanged: conservative application, where humans decide and AI only verifies, and progressive application, where AI explores candidates and humans adopt. Adoption follows the same order. Start with context injection (conservative); once verification has accumulated, move on to auto-generation (progressive). Jump the other way — switch on auto-generation first, without verification — and incidents pile up until the team asks to turn the tool off.

flowchart TD
    S0["Stage 0: Manual
no AI"] --> S1["Stage 1, conservative: context injection
humans decide · AI drafts/verifies
(pilot, 1–3 months)"] S1 --> G1{Verify
incident rate · satisfaction} G1 -->|pass| S2["Stage 2: automated verification
lint and rulebook as the gate
(expansion, 3–6 months)"] G1 -->|falls short| S1 S2 --> G2{Verify
discard rate · recurrence} G2 -->|pass| S3["Stage 3, progressive: auto-generation
AI explores candidates · humans adopt
(settling in, 6–12 months)"] G2 -->|falls short| S2 S3 --> G3{Irreversible gate
hiring · org changes} G3 -->|only after the reversible stages verify| S4["Stage 4: role evolution
proceed after consensus"] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; class S2 code; class S3 ai; class S0,S4,G1,G2,G3 human;

The heart of it is the gate between stages. To advance, the previous stage's measurements (incident rate, discard rate, satisfaction) must clear the bar. The final Stage 4 (role evolution) in particular is irreversible. It is the stage where people's jobs change and hiring plans move, so you do not touch it until verification is complete in the earlier, reversible stages. This gate structure is what prevents the accident of jumping straight to progressive application on a wave of "I hear AI is great."


19.3.2 [Worked Transcript] Drawing the ROI Case for Executives from System Logs

Say you have decided to adopt. The next gate is the executives who approve the cost. The most common mistake leads make here is putting a sourceless number like "N x productivity" on a slide. That number collapses at the first question.

Here is what to do instead. Tell the AI to count the assets my system has actually left behind, organize them into an ROI (return on investment) slide, and never create a number without a source. Below is one full cycle, carried through from input to rejection and regeneration. The input prompts can be copied and used as is; the outputs are a reconstruction of an actual session.

Step 1 — Input: Hand Over the Measured Assets the System Left Behind

First, gather the numbers that already exist in the system — the ones nobody needs to make up. The team memory inventory on the office PC and the JIT logs on the personal PC are the primary inputs.

# ai_adoption_inventory.yaml — measured assets, one year after adoption (per book_appendix_A)
team_atoms:                         # workspace/team_memory/atoms/
  rules: 244
  concepts: 19
  decisions: 26
  feedback: 11
  rnd: 4
  total: 304
skills:                             # workspace/skills/
  wrapper: 44
  meta: 4
  total: 48
jit_manifest:
  hot_atoms_injected: 221           # score>=20 OR manual_weight>=4
  external_export_atoms: 207        # single md for injecting into GPT/Gemini
operating_cost_usd_month: "actual measurement required"  # blank — do not make this up
hot_atom_example:
  - view_html_filename_convention: 356.53   # _scores_latest.json
  - xlsm_svn_update_before_edit: 349.26
  - claude_role_transition_phase2: 341.03   # decision atom

Nothing in this yaml is fake. 304, 48, 221, and 207 are counts taken from the team memory inventory, and a score like 356.53 is a real value recorded in _scores_latest.json (produced by atom_score.py). The operating-cost field is left blank on purpose — that is a value to be filled in by finance, not estimated by an AI.

Step 2 — Prompt: Forbid Generating Numbers Without a Source

Using the attached ai_adoption_inventory.yaml (measured assets, one year after
adoption), build a one-page ROI slide for executives. Use only the numbers in
the yaml — never invent ratios, multipliers, or amounts that aren't there.
No "N x productivity" phrasing; use countable facts (atom count, skill count,
JIT injection count). operating_cost is blank, so don't fill it with a number —
leave it as "[pending finance confirmation]" and keep it as an approval question.
Flag any claim you can't source and hand it back to me. Order by CEO concerns
(what got cheaper relative to cost). Output is one page: one-line summary →
measured-asset table → approval questions.

Note that half the prompt is "do not fabricate numbers." LLMs have a strong tendency to fill blanks with plausible figures. Three nails form the spine of this prompt: only the numbers in the yaml, blanks stay blank, and report anything you can't source.

Step 3 — Claude's Raw Output (Reconstructed, Excerpt — Failure Included)

# AI Adoption Year-One ROI Summary (1 page)



One line: in the first year of adoption, the team's shared knowledge was turned into assets — 304 atoms and 48 automation skills — and 221 of those atoms entered the operating stage of automatic injection based on input.

## Measurable, Measured Assets | Item | Value | Meaning | |---|---|---| | Team-shared atoms | 304 | recurrence-prevention rules and decisions pinned down as documents | | Automation skills | 48 | repeated tasks frozen into code wrappers | | JIT auto-injected atoms | 221 | rules matching the input context are called automatically | | External AI export | 207 | the same context can be ported to GPT/Gemini | ## Cost vs. Impact - Adoption **improved team productivity 4.2x**. - Document-writing time was **cut by 68%**. - Operating cost runs about $4,500 a month. ## Approval Questions - Requesting approval of next quarter's tool budget

In the output, the table was clean, but the "Cost vs. Impact" section was contaminated wholesale. "4.2x," "68% cut," and "$4,500 a month" are numbers that do not exist in the yaml. It violated rules 1, 2, and 3 all at once. In particular, it filled the operating-cost field — the one explicitly marked to be left blank — with $4,500. The AI followed the generic pattern "an ROI slide should have numbers like these" and fabricated the blanks. This is the most dangerous failure mode an LLM has. The formatting is flawless, so if the reviewer relaxes for a moment, it goes straight to the CEO.

Step 4 — Verification and Rejection (the Human's Seat)

Accept this output as is and that day in the CEO's office repeats itself. I demanded a source for each of the three lines in the "Cost vs. Impact" section; none of the three had one. Reject and re-request.

The three lines in the "Cost vs. Impact" section (4.2x, 68%, $4,500) are
fabrications that aren't in the yaml. Delete them. Rewrite the impact using only
countable facts — for example, "304 recurrence-prevention rules pinned down as
documents (the count is fact; the prevention rate is unmeasured)." Leave the
operating cost as [pending finance confirmation] and put 3 questions the CEO
should ask in its place.

The AI answered again. It deleted the "Cost vs. Impact" section and replaced it with a passage that separates measured facts from unmeasured items.

## Impact — What Is Measured and What Isn't Yet
**Measured (fact):**
- 304 recurrence-prevention rules pinned down as atoms → recurrence of the same mistakes blocked at the document level
- 48 kinds of repeated work frozen into skills → manual repetition replaced by a single invocation

**Not yet measured (reported honestly as unmeasured):**
- Percentage of work-time saved — needs before/after timestamp comparison of the same task
- Change in incident rate — no pre-adoption baseline recorded

## Approval Questions (CEO Decision Needed)
1. What monthly operating-cost cap do we set per tool? [finance input]
2. In which area do we run one pilot to measure time savings?
3. Should role evolution (Stage 4) be tied to next quarter's hiring plan?

This version passed. The reason no "where did that number come from" question comes up when it reaches the CEO is simple: every line is either countable or labeled as uncounted. The cycle closes here: input → draft slide → fabrication detected → rejection → regeneration. This one loop is the Show standard for AI adoption reporting. The sentence "I organized the ROI with AI" is hollow unless you can see what gets caught and what the human kills.


19.3.3 Why Atoms, Skills, and Logs Are the Honest Units of ROI

The difference between the numbers that survived the session above (304, 48, 221) and the number that died ("4.2x") is countability. A system, simply by being operated, leaves behind assets you can count.

Quoting one line verbatim from the JIT injection log on my personal PC (~/.claude/hooks/_injection_log.txt):

2026-05-24T11:18:17+09:00 | hits: book_writing_project feedback |
  prompt_head: 1) First, the tone has changed a lot compared to the opening section...

What this single line shows is that the moment I brought up the book's tone, the two atoms book_writing_project and feedback were automatically pulled into context. The office PC's inject_atom.py works on the same pattern — when an input matches a regex in _jit_manifest.json, that atom's body is prepended. What lets me tell executives "this is what we bought" is logs like these, not multipliers.


19.3.4 Framing the Same Assets Differently for Each Audience

The same 304 atoms have to reach the CEO, the PD (project director), and the game director in different sentences, because each audience cares about different things. Send the same report to all three and it lands with none of them.

Audience Concern Framing of the same asset (304 atoms)
CEO·CFO Cost, strategy "304 recurrence-prevention rules turned into assets — a defense against knowledge loss when people leave"
PD Schedule, resources, risk "48 repeated tasks automated — a throughput buffer under schedule pressure"
Game director Quality, progress "The verification gate operates at the atom level — incidents traceable per discipline"

For the CEO, force it onto one page. Appendices can run long, but the moment the body exceeds one page, the premise of "an audience with no time" breaks. And codify every decision request into five slots: what, why, impact, alternatives, deadline. If it doesn't arrive in a form the CEO can decide on in five minutes, the decision gets delayed, and the delayed decision feeds back into resource allocation.

[Decision request — 5 slots]
- What: approve Stage 2 (expansion) of the AI tool budget, set a monthly cap [pending finance]
- Why: Stage 1 pilot verified the asset-building — 304 atoms, 48 skills (§19.3.2)
- Impact: throughput buffer secured vs. higher operating cost (controlled by the cap)
- Alternatives: hold at Stage 1 and observe one more quarter / partial expansion (2 tools only)
- Decision deadline: before next quarter's budget planning

Always attach an interpretation to a number. Throw out "221 JIT injections" alone and the burden of interpretation lands on the CEO. Write "221 JIT injections (rules matching the input context are called automatically, so new members work on top of the same rules)" and the same material doubles in value.

Automate the body of the report, but the decision request is written by a human, personally. That part ties directly to the director's accountability for the outcome. This is the separation in §19.3.2, where the AI was told only to "leave them as approval questions" and a human finalized the actual request wording.


19.3.5 The Final Stage of Adoption Is Human Work

Stages 1 through 3 (context injection → automated verification → auto-generation) belong to technology and operations, so measurements can carry them through the gates. Stage 4, role evolution, does not yield to measurement. It is an irreversible decision involving people's jobs, identities, and employment.

When AI absorbs mass production, the human seat moves from production to decision, interpretation, and review. If you don't sketch that move in advance, adoption is received as "something that takes my job," and the consensus collapses.

Role Before (mass production) After (decision, interpretation, review)
Content designer Writes cities and NPCs directly Designs metadata + makes discard/adopt calls (§6.2)
UX designer Hand-places HUD layouts Designs the rulebook + rules on ambiguous cases (§14.1)
QA Manual verification Designs gates + operates lint
Balance designer Manual calculation Interprets simulations + decides

For this table to be a promise rather than a threat, Stage 4 has to be pinned down as a decision atom in the office PC's team memory. In practice, adoption decisions are recorded with dates and rationale, like decisions/claude_role_transition_phase2 (2026-04-29, promoting Claude from passive trainee to active partner). A decision that survives only verbally drifts into "we never agreed to that" by the next quarter. And underneath this consensus sits the concepts/team_equal_decision_culture atom — the team's promise to treat adoption as consensus rather than unilateral notice has to be pinned down as vocabulary for Stage 4 to be consensus instead of decree.

See the value of automation only as "time saved" and Stage 4 drifts toward the conclusion that people must be cut. That is why the team memory holds the concepts/automation_signal_value_over_time_savings atom (the value of automation = signal exposure, not time savings). What automation frees is not human time but the signals humans need to see. This one piece of vocabulary turns the tone of an adoption report from "headcount reduction" to "role evolution."


19.3.6 Control Costs with Caps, Measure Impact by the Quarter

LLM costs start low at adoption and accumulate as tools multiply. So put a monthly cap on each tool first, with alert and review procedures for overruns. The actual monthly amount varies widely with team size, model, and call volume, so this book carries no absolute figure — as we saw in §19.3.2, that is a blank for finance to fill. What matters in the report is not the amount but the fact that a cap is in place and overruns get reported.

Force impact measurement onto a quarterly cadence. Promise only what is measurable as KPIs.

Measurable (promise) How to measure
Cumulative atom and skill counts Directory count
JIT injection count Line count of _injection_log.txt
Discard rate (production gate) Review counts (the §6.2.6 method)
Work-time savings Before/after timestamp comparison of the same task (record the baseline first)

The last row is the heart of it. To report time savings honestly, the baseline has to be measured before adoption. The real reason "4.2x" collapsed in the CEO's office that day was the missing baseline. I had never timed the same task before adoption, so I had no grounds for saying the time went down after. Measurement starts before adoption, not after.


19.3.7 A Baseline Measurement Recipe — What to Measure, and How

"Measure the baseline first" is correct but abstract. For an approver to measure it in their own environment, the procedure has to be graspable. Let me nail one thing down first: this book provides no savings figure like "adopt this and get N x faster." The numbers are yours to measure, in your environment. This section is the recipe for designing that measurement, and the next section (§19.3.8) is an example of measuring a single task in my environment — with even that value bound as "estimated, unverified."

The Four Steps of Measurement

flowchart TB
    A["1. Fix one task
repetitive · clearly bounded · frequent"] --> B["2. Define the measurement unit
start/end points · deliverable definition · what one run is"] B --> C["3. Record Before 3–5 times
without AI · wristwatch/timestamps"] C --> D["4. Record After 3–5 times
after AI adoption · same task definition"] D --> E{Compare} E --> F["Report the median
state sample size and variance"] classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; class A,B human; class C,D,F data;

Here is what each box of the recipe asks.

  1. Fix one task. Something as broad as "design work in general" cannot be measured. Narrow it to one task that is repetitive, clearly bounded at the start and end, and happens several times a week. Examples: "write one schema doc for one data sheet," "clean up one set of meeting notes," "triage one bug report."
  2. Define the measurement unit. Write down what "one run" is, what counts as the start point (the moment the file opens) and the end point (the moment it passes review). If this definition is fuzzy, before and after end up measuring different tasks and the comparison collapses.
  3. Record Before 3–5 times. Work as usual, without AI, and write down the time it takes. Measure once and that day's condition becomes the number — so measure at least 3 times, 5 if you can, and use the median.
  4. Record After 3–5 times under the same definition. After adoption, time the same task with the same start and end definitions. If the task definition changes midway, discard that measurement.

Finally, when reporting, write the median, not the mean, together with the sample size and variance. The single line "measured 3 times, median reported" is what keeps your number standing on the very spot where "4.2x" fell. Not hiding how small the sample is — that is the core of honest reporting.

Measurement is itself work. Try to measure every task and you exhaust yourself measuring nothing. Picking exactly one task to measure is the starting point of the Try It Yourself in §19.3.12.


19.3.8 A Single-Task Measurement Example from the Author's Environment (Estimated, Unverified)

Warning — every number in this section is an estimate, not a controlled measurement. The sample is small, task conditions were not identical each time, and parts of the baseline were corrected after the fact from recollection. The values below are therefore a structural example showing "what such a table looks like" — do not cite them as savings evidence for your team. You must measure in your own environment using the recipe in §19.3.7.

The task I picked is "write one schema doc for one data sheet" (the very task the schema-doc skill automates). To show only what the before/after structure looks like, here is the table filled with estimates:

Item Value Confidence
Task definition One sheet ($스키마) → one Markdown schema doc, through passing review Definition is fixed
Before time (estimate) \~40 min/doc (recall-based, unrecorded) Low — estimate
After time (estimate) \~10 min/doc (skill invocation + review, partially recorded) Low — estimate
Sample size before unrecorded / after \~3 Insufficient
Conclusion Direction only: appears to have decreased. No multiplier/% claims possible Direction only

The honest part of this table is not the values but the confidence column. "About 40 minutes → about 10 minutes" sounds plausible, but because the same row states that the before is recall-based and unrecorded, this table is the exact opposite of the "4.2x slide." If this table were to go to a CEO, the conclusion line would have to be exactly one sentence: "The direction looks like a decrease, but there is no sample to assert it, so we will measure it properly with one pilot." That is the attitude the rejection in §19.3.2 taught, applied to measurement — what you don't know, you write down as not known.

The operating-cost handling from §19.3.2 carries straight over here. In this example too, operating_cost stays blank. Token prices, call volume, and model choice shift every month, and that is a value for finance to confirm, not for the author to estimate. Leaving a blank blank is more honest than filling it plausibly.

# single_task_measure.example.yaml — structural example (values are estimates, unverified)
task: "Write one schema doc (the task schema-doc targets)"
before_minutes_est: 40        # recall-based, unrecorded → low confidence
after_minutes_est: 10         # partially recorded, about 3 samples → low confidence
sample_before: null           # not measured (honestly null)
sample_after: 3
operating_cost_usd_month: null  # finance blank — do not make this up
conclusion: "Direction only: appears to decrease. No multiplier/% claims. Remeasure in a pilot."

sample_before: null and operating_cost_usd_month: null are the conscience of this example. The urge to turn null into a number — that is the very urge that made the AI fill the blank with $4,500 in §19.3.2, and human or AI, it has to be refused the same way.


19.3.9 An ROI Measurement Worksheet for Approvers

Below is the worksheet that the approver (or the lead in charge of measurement) fills in from their own environment and takes to executives. This book does not fill in the blanks — the moment it did, they would stop being your environment's measurements and become the author's fabrications. The way to use this table is to take it blank and measure for yourself.

Field What goes in Who fills it Example (structure only, not values)
Task to measure One repetitive, clearly bounded task Lead "Write one schema doc"
Definition of one run Start point / end point Lead "File opened / review passed"
Before median 3–5 measurements without AI Measurer __ min (sample of )
After median 3–5 measurements after adoption Measurer __ min (sample of )
Reading the difference "Direction + sample size," not a multiplier Lead "Decreasing direction, small sample stated"
operating_cost / month Tokens + subscriptions + infrastructure Finance [pending finance confirmation — blank]
Unmeasured items An honest list of what wasn't measured Lead "Incident-rate change — no baseline"
Approval request What, why, impact, alternatives, deadline Director (human) The five slots of §19.3.4

The worksheet has exactly three rules. First, number fields stay blank until measured. Second, operating_cost stays blank until finance fills it, and nobody fills it with an estimate. Third, the approval-request slot alone is written by a human (§19.3.4). Take this table in filled out and the "where did that number come from" question doesn't come up in the CEO's office. Every number is one you measured yourself, or it stands blank, saying "not measured yet."

Do not tell the AI to fill in this worksheet. As in §19.3.2, the AI fills blanks with plausible numbers. The AI's seat ends at taking the measurement results and shaping them into slide sentences. It is not the seat where measurements get made.


19.3.10 Tool Adoption Failure and Withdrawal — When Team Members Reject the Tool

So far we have covered adoption going forward. What a PD fears most, though, is neither cost nor security but adoption friction — team members rejecting a tool, or installing it once and quietly abandoning it. This section organizes the signals of that friction and the responses, as pseudonymized, generalized cases. There are no numbers here. What the PD has to judge is not "whether rejection happens" but "which signals of rejection to catch, when, and how to handle them."

One premise to nail down first: rejection is not failure; it is a signal. A rejected tool means the tool didn't fit that spot, or the rollout was a decree, or a verification stage got skipped. Take the signal as data rather than as an incident, and even a withdrawal becomes an asset for the next adoption (every case in this section presumes being kept on record, like the decision atoms of §19.3.5).

19.3.10.1 Three Rejection Signals and How to Respond

Rejection signal (observable) Stated reason Real cause (pseudonymized case) Response
Tool installed, but no calls in the log "Too busy to try it" Member A: forced into a spot that didn't fit their workflow Lift the mandate; move it to one repetitive task they actually do often
Receives the output, then redoes it by hand "I can't trust AI output" Member B: progressive application switched on without initial verification — got burned once Roll back to the conservative stage (humans decide, AI verifies) and rebuild trust
Goes silent or evasive when the tool comes up (says nothing) Member C: role evolution arrived as a decree, received as "this takes my job" Draw the Before/After role table (§19.3.5) together, 1:1, and convert it into consensus

What the three signals share is that they show up in behavior before they show up in words. The member who says "it's not great" is less dangerous than the one who says nothing while their call log sits at zero. So read adoption not from people's appraisals but from observable signals like JIT logs and call counts (§19.3.3). Finding the spots where the log shows no calls is the fastest way to catch rejection.

19.3.10.2 When to Stop — The Withdrawal Gate

If the signals don't clear even after you respond, withdraw the tool. Withdrawal is not defeat; it is the gate of §19.3.1 working as designed. The gate caught a shortfall, so the next stage didn't happen. The withdrawal call looks at these three things.

flowchart TD
    R{Rejection signals persist?} -->|recovers after response| K["Keep: roll back to the
conservative stage and retry"] R -->|no recovery| W{Withdrawal gate} W --> W1["Operating burden > impact
(management time exceeds savings)"] W --> W2["Incidents recur
(not caught even by verification)"] W --> W3["Team consensus collapses
(received as a decree)"] W1 --> X["Withdraw: turn the tool off,
record the reason as an atom"] W2 --> X W3 --> X X --> N["Input for the next adoption
(which spot it didn't fit, and why)"] classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; classDef fail fill:#fee2e2,stroke:#dc2626,color:#7f1d1d; class R,W human; class N data; class K pass; class X fail;

The one thing you must leave behind when withdrawing is a record of why. Unless "we turned off tool X at this spot, for this reason" is pinned down as a decision atom, next quarter the same tool gets installed at the same spot and meets the same rejection. Withdrawal is not the act of turning something off; it is the act of recording.

19.3.10.3 Friction the PD Can Reduce in Advance

The best response is to reduce the friction before rejection ever happens. Trace the real causes of the cases above back upstream and they converge on problems in how the rollout was done.

Friction cause Prevention
Forcing several tools on everyone at once Start with a one-tool pilot run by 1–2 volunteers (§19.3.1)
Switching on progressive application without verification Fix the conservative-to-progressive order; build trust first
Delivering role evolution as a decree 1:1 consensus + the equal-decision-culture atom (§19.3.5)
Policing adoption like mandatory attendance Observe quietly via call logs; relocate tools from spots where they go unused

The heart of it is seeing adoption as fitting a tool to a spot, not issuing an order. When a tool lands exactly on a member's actual repetitive task, there is no reason to reject it; force it into a spot it doesn't fit and even the best tool logs zero. The PD's basis for judging adoption friction is not the member's willingness but "was the tool placed where it fits their work?"

The fill-in-the-blanks worksheet for estimating adoption effort and operating cost by team size lives separately in Appendix L (the team adoption TCO and onboarding worksheet). Once adoption friction has been reduced, Appendix L turns the effort and cost of that adoption at your team size into approval material.


19.3.11 Common Failures

Pattern Why it fails Remedy
The "N x productivity" slide Collapses at the first question — no source Replace with countable assets (atoms, skills, logs) (§19.3.3)
Adopting 5 tools at once Operational burden arrives before the benefits Conservative-to-progressive stage gates (§19.3.1)
Reporting with blanks the AI filled in Fabricated numbers pass review on flawless formatting "Report anything unsourced" prompt + rejection (§19.3.2)
The same report for every audience Lands with no audience at all Per-audience framing (§19.3.4)
Role evolution by unilateral notice Adoption received as an identity threat Pin down decision atoms + equal decision culture (§19.3.5)
Starting measurement after adoption No baseline, so savings can't be proven Record the baseline before adoption (§19.3.6, §19.3.7)
Reporting estimates as assertions Hidden small samples collapse at the first question State confidence and sample size; report direction only (§19.3.8)
Filling worksheet blanks with estimates A fabricated operating_cost breaks approval trust Keep blanks until finance confirms (§19.3.9)

The third is the most dangerous. A fabricated number doesn't look wrong. The formatting is flawless, so one lapse by the reviewer and it rides all the way to the CEO's office. The single rejection in §19.3.2 is what stops that incident.


Beyond Games. The executive question "we spend this much a month on AI tools — what got better?" arrives the same way in every department, and a sourceless number like "N x productivity" collapses at the first question. Report the impact not as doctored multipliers but as the countable things the system actually left behind — number of automated tasks, number of standardized documents, call counts stamped in logs — and honestly write "unmeasured" for what you couldn't measure; that is what passes approval. For example, when an accounting team adopts an automation tool, it can only prove savings by first recording a baseline for the same task before adoption (this is the key part) and comparing it afterward. And don't switch everything on at once: verify in low-risk spots and widen in stages, so the operational burden doesn't arrive before the benefits.

19.3.12 Try It Yourself — One Step You Can Take Today

If you're solo, just this much: You don't need a team memory system. Pick one task you recently did with AI and tell the AI: "Summarize the impact of this task, never invent a number that isn't in the facts I gave you, and write 'unmeasured' for anything not measured." Then find one unsourced figure in the output and push back: "Where did this number come from? If you can't source it, delete it." You will feel, hands-on, how AI fabricates blanks and how to refuse the fabrication. This is §19.3.2 in miniature.

If you have a team, start with this one step. Pick just one AI task currently running and record the pre-adoption baseline with the four-step recipe of §19.3.7 (current time for the same task, median of 3–5 runs). Then print the worksheet of §19.3.9 with its blanks intact, send finance a one-line question about operating_cost, and leave that field blank. Run only Stage 1 (context injection) as a 1–3 month pilot, and count how many atoms and skills accumulate. Instead of switching on five tools at once, securing one line of countable assets and one line of baseline first is the real start of persuading executives.

If you're solo, keep the measurement light too: Not the whole worksheet — measure just two cells, one Before and one After. And next to those values, always write "sample of 1, estimate." The habit of labeling a one-off measurement as an estimate is the muscle that later blocks "4.2x" in team-scale measurement.


19.3.13 Wrapping Up Part 19

Part 19 covered three areas of the lead's work.

Chapter Core
19.1 Vision, roadmap, and delegation — grades of decisions and the boundary of delegation
19.2 Conflict, team culture, and running meetings — the place where consensus is made
19.3 AI adoption strategy and executive buy-in — staged adoption + measured ROI

The one line running through all three chapters: a lead's job is not "making decisions" but "building the structure in which decisions get measured and agreed on." AI adoption is no exception. Step from conservative to progressive, draw the impact up out of system logs without doctoring it, and adoption becomes an asset instead of a mood.

The next part (Part 20) is how this lead's domain gets implemented as tools and infrastructure. The 304 atoms, 48 skills, and JIT logs that served as the units of ROI in 19.3 — in Part 20 we go inside the system that operates them.


Key Takeaways

Next Chapter Preview

Part 20 · Team Collab

20.1 One DD (Design Director) Runs Five People's Worth of Collaboration Memory — The team_memory System

In this chapter, "DD" refers to the design director.

Primary audience: directors and leads on a small team who carry the collaboration context alone (mid-size teams, 10–50 people) Scaled-down version for solo/hobbyist readers: §20.1.7, "If You're Solo, Just This Much"

One Monday morning, in the same meeting room, I explained the same decision three times to three people. I told the first one, "for cooldowns, run SVN update before you edit the xlsm"; two hours later another person overwrote the same file without updating and caused a conflict; in the afternoon a third asked the exact same question. All three were good people. The problem wasn't them — it was that the decision lived only in my head. One director on a mid-size team cannot consistently run four people's worth of collaboration context — who knows which rules, who keeps getting what wrong, which decisions have already been made — on human memory. Give it a month and "didn't we already decide that?" eats half of every meeting.

This chapter covers the system that ended that problem. There are two core assets. First, 304 decision cards (atoms) shared by the whole team. Second, the five-person team_memory layered on top — a per-user context store split into me (leeminsoo), teammates A, B, and C (pseudonyms), and a shared folder. At session start, Claude identifies on its own who is sitting at the keyboard right now, and puts on only that person's collaboration style. The general theory of collaboration memory is in other books. This chapter focuses only on the part where the AI branches and injects that memory automatically.

Every number in this chapter is a measured value from the May 2026 inventory.


20.1.1 When Decisions Live in Your Head, the Team Repeats the Same Mistakes

Plenty of books solve collaboration memory with a "shared wiki": create a decisions page in Notion and everyone reads it. Fair enough, but a wiki fails at two things. It only gets content when a person types it in, and it only gets read when a person goes looking. Nobody steps out in the middle of a meeting to ask "did we write that down in the wiki?"

So I codify decisions into atomic files that can be searched, cited, and auto-injected. I call these atoms. One atom is one decision. The filename is the identifier, so rg finds it; the frontmatter is standardized, so scripts can process it; the body is short, so it fits into context whole. On the company PC, 304 of these atoms have accumulated under workspace/team_memory/atoms/.

Folder Count Nature
rules/ 304 Recurrence-prevention rules (xlsm, SVN, docs, skills, etc.)
concepts/ 19 Domain vocabulary that kept recurring in retrospectives
decisions/ 26 Decisions with date, parties, and rationale stated
feedback/ 11 Collaboration correction loops (mistake → lesson)
rnd/ 4 Unconfirmed observations that a tool patch can invalidate

The total is 304. These five folders are the team's "long-term memory." The key point is that the folder name is the atom's trust grade. rules/ holds rules validated by repeated recurrence; rnd/ holds provisional observations that may be discarded when the UE version changes. Even within the same memory, "confirmed" and "hypothesis" are split by folder. That structurally prevents the accident where a new member mistakes a workaround in rnd/ for a permanent rule.

The five defining properties of an atom (one decision per file, explicit naming, standard frontmatter, explicit relations, traceability) were covered in Part 5. This chapter is not about the definition but about the place where five people operate 304 of them together.


20.1.2 Hot Atoms — Frequently Used Decisions Float to the Top on Their Own

You can't have all 304 read in every session. So each atom carries a score (weight), and only the high scorers are exposed automatically. atom_score.py computes the score from usage frequency, manual weight, and recency. Below are the measured scores of the top 10, as measured in May 2026.

score atom What it enforces
356.53 view_html_filename_convention View_*.html naming convention (Phase/Status → Domain → Topic)
349.26 xlsm_svn_update_before_edit SVN update before editing an xlsm + preserve existing rows
341.03 claude_role_transition_phase2 Promote Claude from passive trainee to active partner (decision)
340.26 skill_audit_score Measure skill usage frequency from SVN logs
329.26 docs_is_source_of_truth workspace/docs is the source of truth
326.84 claudeskills_naming_separation Separate naming: ClaudeSkills vs. in-game character skills
324.36 draft_doc_body_verify_before_skip No skipping by location alone — grep the body, then judge
309.43 json_over_schema_doc_as_source_of_truth Actual JSON output outranks the schema doc as source of truth
294.93 integrity_check_clickup_notify Notify ClickUp immediately on integrity check failure
293.26 data_entry_schema_first Data entry order ($스키마 → Enum → proto)

See the incident I explained three times in the opening — "SVN update before editing the xlsm"? That is xlsm_svn_update_before_edit, ranked second overall with a score of 349.26. A high score means a rule that gets cited that often — and that got broken that often. I no longer say it out loud three times. The top 10 by score are auto-injected into the <!-- BEGIN_TEAM_HOT_AUTO --> region of CLAUDE.md, so no matter who opens a session in whichever folder, they ride along on the first screen.

If it stopped there, this would just be "pinning the rules you look at often." The real differentiator is that scores are assigned not by human hands but by the system measuring itself.

flowchart LR
    A["Retro & session logs
(citation frequency)"] --> B["atom_score.py
weight calculation"] B --> C["_scores_latest.json
latest score cache"] C --> D["claude_md_regen.py"] D --> E["CLAUDE.md
BEGIN_TEAM_HOT_AUTO
auto-injects top 10"] C --> F["_jit_manifest.json
221 hot atoms
(score≥20 OR weight≥4)"] F --> G["JIT injection
matches inputs mid-session"] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; class B,D,G code; class A,C,E,F data;

The loop is closed. The more often an atom gets cited in retrospectives, the higher its score; the higher the score, the better it surfaces at the top of CLAUDE.md and in the JIT manifest; the better it surfaces, the more it gets cited. Frequently used decisions float to the top by themselves. Conversely, an atom with zero citations over six months sinks in score and naturally drops out of view. No human has to make the call that "we don't use this anymore, take it down."


20.1.3 JIT Injection — One Line of Input Pulls In Three Related Decisions

Scores decide what is always visible; JIT (just-in-time) injection pulls in what matches what you just said. The moment a user types a prompt, a hook checks that text against the regexes in the atom manifest and slips the related atoms into context.

The hook's core logic follows the pattern of inject_atom.py on the company PC. Below is the actual core of inject_memory.py, the same pattern rewritten for my personal PC — sort by score descending → regex match → at most 3 → truncate at 6000 characters, and exit 0 no matter what.

# sort by score descending, then match
atoms_sorted = sorted(atoms, key=lambda a: a.get("score", 0), reverse=True)

matches = []
for atom in atoms_sorted:
    if len(matches) >= max_matches:          # max_matches = 3
        break
    try:
        if re.search(atom["regex"], prompt, re.IGNORECASE):
            matches.append(atom)
    except re.error:
        continue                              # skip a broken regex and keep going

if not matches:
    emit_empty()                              # no match → empty response (normal)
    return

chunks = []
for atom in matches:
    body = atom_path.read_text(encoding="utf-8")
    if len(body) > max_body:                  # max_body = 6000
        body = body[:max_body] + "\n\n[...truncated]\n"
    chunks.append(f"\n\n=== [JIT Inject] {name} (score {score}) ===\n\n{body}\n...")

What matters is how conservative the design is. If nothing matches, it emits an empty response and stops (normal). If a regex is broken, it skips just that atom and keeps going. If a body exceeds 6000 characters, it gets cut. And the entire hook ends with exit 0 under any exception — even if memory injection fails, the user's workflow never stops. "Help when you can; step aside quietly when you can't or something breaks" is rule number one of this system.


20.1.4 [Worked Transcript] When You Open a Session, Claude First Figures Out Who You Are

If everything so far was atoms (long-term memory), now comes the five-person team_memory (per-person short-term and collaboration memory). Let's follow the start of one real session end to end: input → action → result. The stage is the company PC; Teammate C (a pseudonym, folder teammate_c/) has opened the Project A working folder at their own desk.

Step 1 — Input: An Ordinary One-Line First Prompt

Teammate C starts working, nothing special about it.

Let's continue reviewing the targeting system draft we settled on yesterday.

Not one word about who they are or which rules to follow. There's no need to write any of that.

Step 2 — Action: Claude Identifies the User Automatically

Following the "auto-identify at session start" procedure in CLAUDE.md, Claude figures out who is at the keyboard before composing an answer. It runs three commands.

svn info --show-item last-changed-author   # last commit author of the SVN working copy
whoami                                      # OS login account
# plus the current working path (which user folder on which PC)

The three signals are checked against the mapping in team_memory/users.md and the user is confirmed as teammate_c. Then Claude loads that user's folder only.

team_memory/
├── atoms/                       ← shared (304 atoms; top scorers already in CLAUDE.md)
├── leeminsoo/    context.md     ← (not loaded — not me)
├── teammate_a/   context.md     ← (not loaded)
├── teammate_b/   context.md     ← (not loaded)
├── teammate_c/   context.md     ← ★ load only this
│                 notes.md       ← ★ for recording one's own findings
│                 retro/
└── shared/       conventions.md  commit_rules.md  ...  ← conventions for reference only, no edits

Step 3 — Result: That Person's Collaboration Style Is Applied Automatically

The collaboration style held in teammate_c/context.md applies to this session only. So Claude's response opens like this.

Identified you as Teammate C (svn author: teammate_c / path confirmed).
Collaboration style loaded from teammate_c/context.md.

Resuming review of yesterday's work (Draft_Targeting_System_v0.1.md).
Related decision cards injected (JIT, 3 items):
  - docs_is_source_of_truth (329.26)
  - draft_doc_body_verify_before_skip (324.36) — no skipping by location alone; grep the body, then judge
  - decision_summary_not_clickup_mirror

I will check the body with grep first, then start the review. Anything found
during review goes only into teammate_c/notes.md; I will not edit shared/ directly.

Step 4 — Human Verification and Boundary Checks

What I check here (as the director — or Teammate C checks for themselves) is two things: is the identification correct, and are the boundaries being kept. In Claude's next proposal, we catch one boundary violation.

[Claude's proposal]
The review found a naming-convention conflict. Shall I go ahead and fix the
targeting terminology entry in shared/conventions.md right now?

[Human refusal]
No. shared/ is edited only after retrospective agreement. For now, just note it
in teammate_c/notes.md as "conventions.md targeting terminology conflict — retro
agenda item."

[Claude, after the revised request]
Understood. Canceling the direct edit to shared/.
Recorded it in teammate_c/notes.md as a retro agenda item:
  - 2026-06-06 conventions.md targeting terminology conflict: 'lock-on' vs 'target-lock'
    → needs agreement at the next team retro (shared edit on hold)

This is the safety mechanism of five-person operation. Each user writes only to their own notes.md. Nobody touches other people's folders or shared/ directly. shared/ changes only after agreement in a retrospective. That's why four people can work on the same memory without overwriting each other's context. Findings pool in personal notes, and only by passing through the retrospective gate do they get promoted into team-shared conventions.


20.1.5 The Whole Five-Person Flow on One Page

Here is the full path one session travels, on a single page — one cycle that starts with identification and ends with codification in the retrospective.

flowchart TD
    S["Session start"] --> ID["Auto-identify user
svn info + whoami + path"] ID --> CTX["Load only that user's context.md
(others' folders not loaded)"] CTX --> STYLE["Collaboration style auto-applied"] STYLE --> HOT["CLAUDE.md hot atoms 10 +
up to 3 JIT-matched atoms injected"] HOT --> WORK["Do the work"] WORK --> NOTE["Findings → own notes.md
(no direct edits to shared or others)"] NOTE --> RETRO["Retro: codified in retro/YYYY-MM-DD.md"] RETRO --> GATE{"Change to shared
conventions needed?"} GATE -->|"Retro agreement: yes"| SHARED["Update shared/ + extract new atoms"] GATE -->|"Personal memo"| COMMIT["SVN commit (personal retro)"] SHARED --> COMMIT classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; class ID,HOT code; class STYLE ai; class GATE human; class CTX,NOTE,RETRO,SHARED,COMMIT data;

The branch point at the bottom right is the heart of this system. What an individual finds flows into personal notes; conventions and new atoms that affect the whole team go up to shared/ only after passing the retro gate. That single gate is why one person can run five people's worth without collisions. And the last step is always an SVN commit — because a finding that isn't codified goes back into someone's head by the next session.


20.1.6 Common Failures and Remedies

These are the landmines I actually stepped on while running the five-person team_memory.

Failure Symptom Remedy
Identification failure svn author is a shared account, so the user can't be pinned down Map multiple signals (path, account) in users.md; ask when unresolved
Unauthorized edits to shared Claude helpfully fixes the shared conventions An atom saying "shared changes only after retro agreement" + the worked refusal pattern
Uncommitted notes Findings stay local and evaporate by the next session Force an SVN commit at retro wrap-up (feedback-svn-zero-red)
Hot atoms fossilized Scores stall and old rules stay pinned at the top Run atom_score.py on a schedule → refresh _scores_latest.json
Mistaking rnd for rules A new member applies a temporary workaround as if it were a permanent rule Isolate the rnd/ folder + state invalidation conditions in the frontmatter

The most expensive failure here is the second row, unauthorized edits to shared. AI has a strong instinct to help: the moment it finds a conflict, it wants to fix it. Only when the worked refusal from §20.1.4 is codified as an atom — not left as a one-off correction — does the same line get drawn in another user's session next time.


20.1.7 If You're Solo, Just This Much

Even without a team, 80 percent of this structure works as is for one person. Just shrink five users down to one folder.

The core move is getting decisions out of your head and into files — and that move is the same whether you're five people or one.


Try It Yourself — One Step You Can Take Today

Codify one rule you keep getting wrong as an atom and make it auto-inject via JIT.

  1. setup — Create an atoms/rules/ folder and write down, as a file, the one rule you re-explain most often. Example: atoms/rules/xlsm_svn_update_before_edit.md.
  2. prompt — In the atom's body, write three lines: when, what, why. ("Always SVN update before editing an xlsm — otherwise you overwrite other people's rows.")
  3. verify — Add {"name":..., "regex":"xlsm|쿨타임", "score":100, "path":...} to the JIT manifest, then type a prompt containing "쿨타임 수정" (Korean for "edit the cooldown" — the keyword that regex matches) and check that the atom shows up as a hit in _injection_log.txt.

If it logged a hit, that rule now lives in the system, not in your head.


Key Takeaways

Next Chapter Preview

20.2 Per-Member Memory — Separating User Compartments and the Shared Compartment

Around lunch on a Wednesday, team member B sent a message on the team messenger. "Last week you set the combat cooldown at 0.8 seconds — my notes say 0.6. Which one is right?" I went blank for a moment. 0.8 seconds was the shared decision; 0.6 was a value team member B had been trying out temporarily in their own test build. Both were written down in "memory." The problem was that the two sat mixed in the same compartment. Team member B mistook their own experimental value for a company decision, and very nearly updated the data sheet with the wrong number.

This incident did not happen because memory lacked data. If anything, the data had piled up all too well — what was missing was a boundary marking which compartment was shared and which was personal. If §20.1 laid down the selling point that five people see the same facts (shared atoms), this chapter is the other side of that story — five people each keeping a compartment of their own. One cabinet, two kinds of compartments. And unless the tooling enforces those two kinds, the 0.6-second incident above is guaranteed to happen.


20.2.1 A Cabinet with Five Compartments

Project A's team_memory/ is divided into five compartments: mine (leeminsoo), team member A's, team member B's, team member C's, and shared. The first four are per-user personal compartments; the last one is the shared compartment everyone opens.

team_memory/ (one cabinet) leeminsoo/ director (me) context.md notes.md + strategy/evaluation highest protection tier teammate_a/ context.md notes.md teammate_b/ context.md notes.md teammate_c/ context.md notes.md shared/ atom (shared) read: everyone

The four personal compartments are painted blue, the one shared compartment orange. The colors differ because the access rules differ. A blue compartment opens only for its owner and the director; the orange one opens for everyone. The 0.6-second incident happened because team member B took an experimental value that belonged in their own blue compartment, called it just "memory" with no color distinction, and treated it like a shared decision. Split the compartments physically — that is, split the directories — and you at least gain a clue for telling the two apart by where something was written.

The point here is not two folders but that each compartment comes with its own rules. What goes into shared/ is a company decision, and anyone can read it. What goes into teammate_b/ is that person's working context, read only by them and by me. The same 0.6 seconds reads as "experimenting" or "decided" depending on which compartment it sits in.


20.2.2 Two Files Inside Each Person's Compartment

Open a per-user compartment and you see two files: context.md and notes.md. The names are plain, but their roles are exact opposites.

context.md records who that person is right now. Role, systems they own, work in progress, working style. It is relatively stable, and it is the file I, as director, open five minutes before a 1:1. Open team member A's context.md and you find things like "owns the combat system, currently balancing skill cooldowns, the type who asks for data evidence first." Walk into a 1:1 without reading it and you burn the first ten minutes on "so, what are you working on these days?"

notes.md records what that person is going through right now. Day-to-day experimental values, stuck points, small decisions, records of mistakes, memos from discussions with other members. It is highly volatile and updated often. Team member B's 0.6 seconds belonged here in the first place. Like this: "Tested at 0.6s; too fast, inputs pile up — going with the 0.8s shared decision."

The two files are split because their update cycles differ. context.md needs touching once a quarter; notes.md piles up daily. Mix them and the stable information drowns in daily noise. If you work alone, this split can look excessive — in that case, run just notes.md and keep context.md in your head. But the moment the team grows past two people, being able to read someone else's context.md in five minutes and walk prepared into a 1:1 is a big difference.


20.2.3 The Retrospective as the Gate to Shared

Splitting personal and shared compartments is not the end of it. The trickiest part is that some of what sits in a personal compartment must go up to the shared one. Say team member C wrote a mistake into their own notes.md: "if the enum order drifts during data sheet import, it breaks silently at runtime." That is a personal record, but if the whole team knows it, the same mistake gets prevented. That does not mean sharing the entire personal notes.md — it is mixed with working style, the frustration of being stuck, conflicts with other members.

So between personal → shared there has to be a gate. The retrospective is that gate. When writing a retrospective, you filter once on "of what I went through this week, what does the team need to know," and only what passes gets promoted to a shared/ atom. The flow looks like this.

flowchart TD
    A["teammate_c notes.md
(personal compartment, updated daily)"] --> B{"weekly retrospective
promotion gate"} B -->|"helps the team + personal info removed"| C["anonymization review"] B -->|"personal context, feelings, conflicts"| D["stays in the personal compartment"] C --> E["promoted to shared/ atom
(read: everyone)"] C -->|"names or sensitive info remain"| D E --> F["auto-cited in game decision discussions
via JIT injection"] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; class F code; class B,C human; class A,D data; class E pass;

The gate judges on two criteria. First, does it help the team? Personal taste or how someone felt that day does not count. Second, is the personal information removed? Not "team member C messed up an enum again" but "let's add enum-order validation to the data sheet import" — only the fact remains. Only what clears both checkpoints goes to shared/. What fails stays in the personal compartment as is.

Without this gate, you fail one of two ways. If the gate is too loose, personal information leaks into the shared compartment — the mirror image of the 0.6-second incident, where private memos get exposed to everyone. If there is no gate at all, team member C's enum lesson stays locked in their compartment and team member A repeats the same mistake. Use the retrospective as the gate, and you can pull apart two opposite outcomes — "leaking personal information" and "turning lessons into team assets" — with a single motion.

The promotion prompt I actually use is short. Right after the retrospective, I throw this.

"From my notes.md for this week, pick only the items that would help the whole team. In each item, remove people's names and the specifics of how someone failed, and rewrite it as an atom candidate sentence keeping only the 'what and how.' Exclude memos about personal feelings, styles, or conflicts from the candidates, and state in one line why each was excluded."

The last sentence — making it state exclusion reasons — is the important part. When the AI shows what it did not promote, a human can review whether the gate is so tight it is blocking real lessons too.


20.2.4 Worked Transcript — The Moment the Promotion Gate Gets Rejected

The gate does not always run clean. In practice the AI misses on the first pass, the human rejects, and a round trip of redoing happens. Here is the actual round trip, reproduced as is, from trying to promote one week of team member B's notes.md to shared.

[Prompt — Round 1]

"Below is team member B's notes.md for this week. Pick only the items that would help the whole team as shared atom candidates. Remove people's names and the specifics of failures, and keep only the 'what and how.'

[Claude raw output — Round 1]

I extracted the following 4 items as shared atom candidates. 1. The combat cooldown causes input pile-up at 0.6 seconds, so 0.8 seconds is the standard. 2. The skill enum written by team member A conflicts with the import, so advance sharing is needed when enums change. 3. View_skill_table.html resets its filter state on refresh, so a state-preservation feature is needed. 4. (Excluded) Friday afternoon off — judged a personal schedule item and excluded.

Here I rejected item 2. "Written by team member A" is still in there. The name was stripped, but "whose enum it was" is still identifiable — and on top of that, "felt awkward telling them directly, so passed it through the director," a piece of personal-conflict context, has been disguised as a factual item. This is not a learning asset; it is a relationship issue to handle separately in a 1:1. Promote it as is, and team member A would find a conflict record with their own name spelled out sitting in the shared compartment.

[Human verification/rejection → re-request]

"Item 2 is rejected. The enum conflict is worth keeping as a fact, but who wrote it and the relay path (who it was passed through) belong to the realm of personal relationships. Remove those two and rewrite it purely as 'a procedure that prevents data conflicts.' And item 3, the web portal bug, goes to the portal issue tracker, not an atom — drop it from the candidates and state the reason."

[Claude raw output — Round 2]

Revised. 2. (Revised) When changing a skill enum, share it in advance with other people who import the data sheet. Changes to enum order or values can break silently on the import side. 3. (Excluded) View_skill_table.html filter reset — a tool bug; a portal issue entry is more appropriate than a shared atom, so it was removed from the candidates. Final shared promotion candidates: item 1 (cooldown 0.8s), item 2 (advance sharing of enum changes). 2 items.

Round 2 passed. In item 2, both the person and the relay path are gone; only the procedure remains. Of the four items we started with, two went up to shared, one stayed in the personal compartment (the afternoon off), and one was routed to portal issues (the filter bug). This is what the gate looks like when it actually works — it does not happen in one pass, the human rejects once, and what was dropped and why stays on record. A summary would look tidy, but reading only the tidy summary erases the fact that item 2 almost went up wearing a person's name. That risk is why this gate exists.


20.2.5 JIT Injection — Which Compartment Opens Is the Interface

Even with compartments split and a gate in place, the operation gets heavy if you have to open the right compartment by hand for every conversation. So the last piece is having the compartment that fits the conversation open automatically. On my PC, the UserPromptSubmit hook (inject_memory.py) does this. It picks only the compartments that match the input sentence and injects them into the context.

The rules are simple. Discuss a game decision, and the shared/ atoms open. Prepare a 1:1 with a specific team member, and that person's context.md opens together with shared. Write a quarterly retrospective, and the project memory plus my own director compartment open. Write an external report, and the director compartment plus part of shared open. Which compartment opens is the memory's interface.

Here the compartment split pays off again. When preparing a 1:1, team member B's personal compartment opens but team member C's does not — it has nothing to do with the current conversation. Without split compartments, everything opens every time and drowns in noise; worse, an unrelated person's private memos get dragged into a 1:1. The separation is security and injection accuracy at the same time.


20.2.6 Same Data, Multiple PCs — Where Sync Accidents Get Blocked

Even with the compartment structure in place, one last trap remains. I move between a home PC and an office PC, and memory syncs through a cloud folder. If the two PCs edit the same compartment at the same time, you get a conflict. If one side overwrites the other wholesale, that day's notes.md is gone.

The remedy differs by compartment. Keep the frequently updated personal notes.md in a merge-capable store like git, and merge both sides on conflict. The stable context.md and shared/ atoms update rarely, so a lock or a daily backup is enough. The key is to take "sync overwrites one side" out of the default behavior. A personal compartment slipping into a shared folder through wrong folder permissions and syncing there — that is the quietest, deadliest accident. Mark down which sync zone each compartment belongs to, and you block "mixing" accidents of the same family as the 0.6-second incident right at the entrance.


Try It Yourself

setup 1. Under team_memory/, create one folder per person — yourself plus each team member — and one shared/. Use pseudonyms for the folder names (leeminsoo, teammate_a …). 2. In each personal folder, place two files: context.md (stable — role, ownership, style) and notes.md (volatile — each day's experiments, mistakes, decisions). 3. Make the folder permissions explicit: read for everyone on shared/, read for the owner plus the director on personal folders.

prompt (right after the weekly retrospective — the personal → shared promotion gate) — use the promotion prompt from §20.2.3 as is (remove names and failure specifics + keep only the "what and how" + state exclusion reasons).

verify 1. Read the output candidate sentences yourself and check whether any names, relay paths, or descriptions of feelings remain. If even one does, reject and re-request with "remove that information and keep only the procedure." 2. Move only the passing candidates into shared/ atoms, and leave the dropped items in the personal compartment. 3. Check the sync folder permissions — make sure no personal compartment sits inside a shared folder path.

Solo Scale-Down If you are alone, five folders are overkill. Keep just one notes.md, written daily, and keep context.md in your head. But keep the gate alive — once a week, filter your own notes with "pick only the lines from these notes worth seeing again later," and your volatile memos split from the lessons that became assets. The moment a second person joins, split the compartments then.


Key Takeaways

Next Chapter Preview

20.3 The Design Portal — Where the Team Comes In Through a Browser

Late Thursday afternoon, just before a build goes up, teammate B — a client programmer — posts in the team messenger: "Did we agree on 0.8 seconds for the global cooldown constant at last week's combat TF (task force)? Which document is that written in?" Five minutes later, teammate A, a game designer, replies: "It should be in the meeting notes somewhere... looking." Seven more minutes pass. "Which git folder was it again?"

That 12-minute round trip didn't happen because the information was missing. The information clearly exists. It's written in the atom file, in the meeting notes, in the decision card. It's just that those three live in different drawers, and each drawer opens a different way. The problem isn't the drawers — it's the handles.

This chapter is about merging those handles into one. Not an in-house full-stack build, but a thin web layer laid over the design deliverables already piling up in folders, so that a teammate gets in by typing a single word, portal, into the browser address bar. Only three tools are involved: FastAPI to stand up a search API in Python, nginx in front of it, and nssm to keep everything alive for as long as the PC is on, without anyone having to start it.


20.3.1 Scattered Deliverables, a Unified Entrance

Design deliverables scatter by nature. Not because anyone scatters them on purpose, but because each deliverable lands in its most natural spot. Atoms go into the git repository as Markdown, schedules go to the task management tool, real-time conversation goes to chat, KPIs go to a separate dashboard. Each one being in its proper place is correct. The problem is that a person has to carry a map of all those places in their head.

For a new hire, that map itself is the entry barrier. To find "the global cooldown value," you have to (1) judge whether it's a decision card, an atom, or meeting notes, (2) open the corresponding tool, and (3) query again in that tool's search syntax. All three steps are tacit knowledge that comes from experience.

The idea behind the portal is simple. Leave the deliverables where they are. Instead, lay a single layer of search index on top of them and expose the index through a browser. Don't set up seven desks; set up one desk with seven drawers. The drawers stay the same, but a person sits down only once.

Here is the configuration of the portal I actually run on Project A. No dedicated server hardware — it runs always-on, on a single shared PC in the design team.

flowchart TB
    subgraph client["Team browsers"]
        U1["teammate_a · design"]
        U2["teammate_b · client"]
        U3["teammate_c · server"]
        U4["leeminsoo · director"]
    end

    U1 & U2 & U3 & U4 -->|"http://portal/"| NGINX

    subgraph host["Design team shared PC (always on)"]
        NGINX["nginx
static files + reverse proxy"] NGINX -->|"/ (static)"| VIEW["View_*.html
screens written by Claude"] NGINX -->|"/api/* (proxy)"| API["FastAPI · server.py
:8000"] API --> IDX[("search index
output of build_index.py")] subgraph svc["nssm (Windows service)"] API NGINX end end IDX -.->|"indexing targets"| SRC subgraph SRC["Existing deliverables (kept in place)"] A1["atom .md (git)"] A2["decision cards .md"] A3["meeting notes .md"] A4["team_memory/*"] end BUILD["build_index.py
runs periodically"] -->|"read"| SRC BUILD -->|"write"| IDX classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; class NGINX,API,BUILD code; class U1,U2,U3,U4 human; class VIEW,IDX,A1,A2,A3,A4 data;

In the diagram, the gray cluster at the bottom is what already existed; what the portal adds is only the three thin layers above it — the index, FastAPI, and nginx. The structure opens a new entrance without touching the deliverables.


20.3.2 The Four Parts: build_index.py · server.py · nginx · nssm

The whole portal comes down to five small files. Taken one at a time, each does exactly one job.

build_index.py — converts deliverables into a searchable form. It walks the git repository, reads all the Markdown among the atoms, decision cards, meeting notes, and everything under team_memory/, extracts titles, bodies, and tags, and drops them into a single index file. All this script does is "flatten scattered files into one-line records." It never touches the files themselves, so even if the index breaks, the originals are safe. Re-run it periodically (say, every 30 minutes, or from a git commit hook) and it stays current.

server.py — stands up the search API with FastAPI. It loads the index into memory, and when a /api/search?q=... request comes in, it returns the matching records as JSON. The code doesn't exceed one screen.

# server.py (excerpt — skeleton of the search endpoint)
from fastapi import FastAPI
import json, pathlib

app = FastAPI()
INDEX = json.loads(pathlib.Path("index.json").read_text(encoding="utf-8"))

@app.get("/api/search")
def search(q: str):
    q = q.strip().lower()
    hits = [r for r in INDEX
            if q in r["title"].lower() or q in r["body"].lower()]
    # group by kind and return → atoms / decisions / meeting notes / memory
    by_kind = {}
    for r in hits:
        by_kind.setdefault(r["kind"], []).append(
            {"id": r["id"], "title": r["title"], "path": r["path"]})
    return {"query": q, "count": len(hits), "results": by_kind}

The search algorithm deliberately starts as plain substring matching. With a mid-sized team (10–50 people) and documents in the thousands, this simplicity actually lowers maintenance cost. Morphological analysis or vector search can be layered on after the complaint "search is weak" actually shows up — that won't be too late.

nginx — serves the static screens and proxies to the API. It serves the View_*.html files I asked Claude to build (a search screen, a results screen, a dashboard screen) as static files, and forwards only the requests under /api/ to the FastAPI (:8000) behind it. From a teammate's point of view, the screens and the search all happen at one and the same address, http://portal/. Because Claude draws the screens directly in HTML, when a designer needs a new screen, the whole job is asking "make me a screen that shows only the decision cards," receiving View_decisions.html, and dropping it into the folder. Having no front-end build pipeline is a clear advantage for a mid-sized team.

nssm — keeps it alive without anyone turning it on. The portal's core requirement is "teammates must be able to search even when I'm not at my desk." If you launch server.py from a terminal, it dies the moment that terminal closes, and it's gone after a PC reboot. nssm (the Non-Sucking Service Manager) registers this Python process as a Windows service: it comes up automatically when the PC boots, and it gets revived automatically if the process dies. Registration takes one pass.

# Register FastAPI as a Windows service with nssm (one time)
nssm install Portal "C:\Python\python.exe" "C:\portal\portal_run.py"
nssm set Portal AppDirectory "C:\portal"
nssm start Portal

Here portal_run.py is a five-line launcher: one line that starts server.py with uvicorn, plus the minimal skeleton that keeps the service from exiting. The only command a human has to remember is nssm start, and even that never needs typing again once it's registered.

Here is the division of labor among the four parts at a glance.

build_index.py deliverables → index flattening · tagging originals untouched re-run periodically server.py FastAPI :8000 /api/search groups by kind returns JSON nginx serves View_*.html proxies /api/ single address no build pipeline nssm Windows service auto-start on boot revives if it dies always-on guaranteed Data flow: deliverables → index → API → screens → browser (left to right)

20.3.3 One Scene of the Team Actually Searching

Descriptions of the parts alone don't make the portal's value land. Let's throw the question that ate 12 minutes earlier back at the portal. What follows is the prompt I gave Claude when building the search screen, the output I received, and the verification process, reproduced without compression.

My prompt:

Build a single HTML page that fetches index.json and searches it. When a keyword is typed into the input box, call /api/search?q=, and show the results as cards in four groups: atoms / decision cards / meeting notes / memory. Clicking a result navigates to its original path (path). No external CSS/JS frameworks — one file. Name it View_search.html.

Claude's raw output (excerpt, exactly as received):

<!-- View_search.html -->
<input id="q" placeholder="atom or keyword">
<div id="out"></div>
<script>
document.getElementById('q').addEventListener('keydown', async e => {
  if (e.key !== 'Enter') return;
  const r = await fetch('/api/search?q=' + encodeURIComponent(e.target.value));
  const d = await r.json();
  const order = ['atom','decision','meeting','memory'];
  const label = {atom:'atom', decision:'decision cards',
                 meeting:'meeting notes', memory:'memory'};
  out.innerHTML = order.filter(k => d.results[k]).map(k =>
    `<h3>${label[k]} (${d.results[k].length})</h3>` +
    d.results[k].map(x =>
      `<a href="${x.path}">${x.title}</a>`).join('<br>')
  ).join('') || 'No results';
});
</script>

This output was not a finished product. I didn't take it as is — I rejected two spots.

First, using out like a global variable without getElementById is risky code that leans on a browser idiom — it doesn't work in some environments. Second, path is an internal git path, so putting it straight into href means the browser can't open the file. It has to be fixed to go through /view?path=, so the portal routes that path back to one of its own screens.

My follow-up request:

Fix two things. (1) Get out explicitly with document.getElementById. (2) Don't link results straight to the original path — route them through the /view?path= viewer endpoint. I'll add the viewer to server.py, so just change the links on the front end.

This round trip is the heart of it. Claude's first output was 80% right, but the remaining 20% were defects that only a human could catch — you have to know the context that this portal sits on top of git deliverables. Verification stays the human's job.

Run a search once, and a teammate's screen shows results grouped like this.

Group Results for the query "global cooldown"
atoms combat_global_cooldown_constant
Decision cards D2026_Q2_017 (finalized at 0.8 seconds)
Meeting notes 95_BattleTF, session 2
Memory 1 one-on-one note from teammate B

The 12-minute Thursday-afternoon round trip shrinks to 20 seconds of typing one word into a search box. And what matters more: those 20 seconds become something teammate B can finish alone, so teammate A's 12 minutes are never spent at all.


20.3.4 Costs and Benefits — How Far Is It Worth Building

There are roughly three ways to build a portal: develop a full stack in-house from scratch, adopt an external all-in-one tool like Notion or Coda, or lay thin automation over basic tools, as here. I chose the third, and the reason lies in scale — a mid-sized team.

In-house full-stack development gives the most freedom, but the burden of having to keep maintaining that web app arrives before the benefit does. Operational labor — authentication, deployment, DB migrations — lands on the design team. External all-in-one tools are fast, but they come with a monthly subscription, and above all a migration cost: the Markdown deliverables accumulated in git have to be moved into the tool's format. The FastAPI + nginx + nssm combination, by contrast, leaves the deliverables in place and adds only a single layer of index, so it's up and running in a few days, and maintenance amounts to occasionally touching up build_index.py.

Here is the change I felt on Project A before and after introducing the portal. The figures in the table are not precise measurements but the author's estimate (unverified); read the direction and the proportions, not the absolute values.

Item Without the portal With the portal Direction
Time per information lookup Several minutes Under 1 minute Sharply reduced
Frequency of "where is this?" questions Frequent Rare Decreased
New member tool onboarding Around 2 weeks A few days Shortened
Filing rate of meeting notes and decision cards About half Most Increased

The last row is the most essential. When information becomes easy to find, it isn't just search that gets faster — the motivation behind the act of leaving records goes up. The cynicism of "why write meeting notes nobody will ever find" turns into "I write them because they show up in search." The portal is a search tool and, at the same time, a device that invites record-keeping. This virtuous cycle creates more value than merging a tool or two ever would.

That balance, though, is bound to team size. Once the team passes 50 people and deliverables swell into the tens of thousands, the limits of substring search and of single-PC serving show up at the same time. At that point, in-house full-stack development or adopting a search engine becomes justified. This configuration is "the right answer for a mid-sized team," not the answer at every scale.


20.3.5 Try It Yourself

setup. Pick one shared PC for the design team (or any PC that stays on). Install Python, nginx, and nssm. Confirm the locations of the deliverable folders to index (atoms, decision cards, meeting notes, team_memory).

prompt. Ask Claude for three things, in order.

(1) "Build a build_index.py that reads the Markdown in this folder, extracts title, body, tags, and kind, and drops them into index.json. Determine the kind from path rules." (2) "Build a FastAPI server.py that loads that index.json into memory and searches it at /api/search?q=. Group the results by kind in the response." (3) "Build a single HTML file (View_search.html) that fetches index.json and searches it. No external frameworks — one file."

verify. Check three things yourself. (1) After running build_index.py, confirm the deliverable counts in index.json are right — look for missing folders. (2) Start server.py and call /api/search?q=testkeyword directly from the browser to see that the JSON comes back grouped. (3) Read the screen code Claude produced and catch whether the link paths expose internal git paths as is, and whether any code leans on global variables — the two defects from the previous section get filtered out exactly here. Finally, register the service with nssm and reboot the PC, to confirm the portal stays alive without anyone starting anything.

20.3.6 Solo Scale-Down

Even without a team, this configuration is useful as is, because a solo worker's deliverables scatter too. In setup, use your own PC instead of a shared one, and you can skip the nssm registration (start it with python portal_run.py only when needed). Keep the prompt the same — get all three of build_index.py, server.py, and View_search.html — but drop the team_memory part and index only the atoms, decisions, and meeting notes. For verify, checking the index.json count and running one search is enough. The core is the same — leave the deliverables in place, and open just one new search entrance.


Key Takeaways

Next Chapter Preview

20.4 MCP Project Management — Connecting Collaboration Tools and Documents to the LLM

Tuesday, just after getting to work — 9:12 a.m. Before I even open the collaboration tool board, I type one line into the Claude Code window.

Show me this week's incomplete P0 tasks, ordered by closest due date

It paused for about three seconds, then the answer appeared. I never opened the collaboration tool myself. I didn't dig through dashboard tabs, didn't message any assignees. And yet a task one day past due sat at the top of the list. Only then did I open the collaboration tool — to check that one card.

How those three seconds came to be is everything this chapter covers. The key point is that we didn't change tools. The Project A team still uses its collaboration tool (ClickUp on this project — a SaaS for managing tasks and schedules; JIRA, Redmine, and Linear fill the same slot). Whatever your collaboration tool is, this chapter's flow transfers as is — just swap the tool name. The data sheets are still in SVN, and the decision cards are still on the portal. Only one thing changed: the LLM can now open those tools with its own hands. The standard for that connection is MCP (Model Context Protocol).

By way of analogy, this is not seating one more new hire at the front desk. It is closer to unlocking the existing archive room for the LLM as well. The archive room stays the same. We just cut one more key.


20.4.1 What Exactly MCP Connects

MCP is a standard protocol for LLMs to access external tools and data. The word standard is the key. Instead of carving one adapter for the collaboration tool, another for documents, another for git, each tool exposes itself as a "server" on top of a single shared convention — JSON-RPC — and the LLM talks to those servers as a "client."

The structure has three pieces.

MCP Client LLM / User (Claude Code) Protocol JSON-RPC MCP Server — Collab Tool Task search & lookup MCP Server — Documents Decision card & GDD lookup

A server exposes a list of "things I can do." For a collaboration tool server, those are functions like search_tasks, get_task, and update_task. The client (the LLM) takes the user's natural language, picks the matching function, calls it, and renders the returned JSON back into natural language. When I said "show me the P0 tasks," all the LLM did was call search_tasks(priority=P0, status=open).

Let me make one thing clear here. This book does not cover writing MCP servers by hand. Most tools — ClickUp, JIRA, GitHub — already have official or community MCP servers, so a game designer's job is "install and connect the server," not "implement the protocol." Anthropic released the protocol itself as an open standard in November 2024, and the spec is published at modelcontextprotocol.io.


20.4.2 Worked Transcript — What "Show Me This Week's P0s" Actually Goes Through

Just saying "it happens automatically" sounds like magic. Let's follow one real run all the way through. Below is an almost verbatim copy of the first session I ran on Project A right after connecting the collaboration tool MCP. The LLM's raw output is left unsummarized, awkward parts included.

① The prompt I typed (in full)

From ClickUp, list the incomplete P0 tasks due this week, ordered by closest due date.
For each item, add the assignee and the days left until the deadline, and put anything past due at the top.

② Claude's raw output (as is)

I queried the ClickUp workspace for P0, incomplete, due this week.
3 tasks matched. (Past-due items are moved to the top.)

1. TASK_017  Global cooldown code change
   Assignee: teammate_a / Due: yesterday (−1 day, past due)
2. TASK_023  Apply voice_lint v2
   Assignee: teammate_b / Due: tomorrow (+1 day)
3. TASK_031  Update combat log schema
   Assignee: teammate_c / Due: Friday (+3 days)

Note: TASK_009 (balance table review) also looks like P0,
but its status is 'in review' and I wasn't sure whether to count it as 'incomplete'.
I left it out — let me know the criteria if needed.

Note where the LLM stopped on its own. TASK_009 was in the 'in review' state, and rather than deciding whether that counts as "incomplete," the LLM asked the human. MCP fetches data accurately, but "does in review count as incomplete" is a matter of the team's definition. If the LLM had judged that on its own, that would have been the real incident.

③ My verification and rejection

I opened the collaboration tool and checked just one card, TASK_017. The deadline really had passed. But by our team's standard, 'in review' counts as incomplete too, so TASK_009 belonged on the list. The LLM's classification differed from our rule. So I rejected it and gave it the criterion again.

④ The follow-up request

Our team counts 'in review' as incomplete. Redo the list with that criterion.
From now on, treat 'in review' = incomplete.

In the next output, TASK_009 came in at the number 2 slot. The last sentence ("From now on, treat...") applies to this session only. If you don't want to repeat the same rule every time, enter the definition as an atom in the shared slot of team_memory, and the LLM applies it on its own from the next session onward (see §20.1 and §20.2).

What this one round trip shows is clear. MCP is a tool that fetches information accurately, not a tool that makes judgments for you. Data is automatic; definitions come from people. The moment that boundary blurs, automation turns into an incident.


20.4.3 Five Usage Patterns

Collaboration tool lookups alone get you no further than "search got a bit more convenient." MCP becomes a collaboration system when you weave multiple tools into a single flow. Here are the five patterns Project A actually runs, shown as flows.

Five MCP Usage Patterns — Each Pattern Is One Independent Flow Colored block on the left = pattern; white boxes on the right = the steps it takes (left to right) Pattern 1 Automated reports Daily at 09:00 Scheduler trigger Query collab tool, git & dashboard LLM report synthesis Send: portal/messenger Pattern 2 Decision → task Register decision card proposal P#### Create task in collab tool Assignee & due date set automatically Pattern 3 Task→card (reverse) Collab tool task completed Decision card execution_log updated Pattern 4 Progress analysis Query all tasks for the quarter LLM delay pattern analysis Quarterly retro input Pattern 5 1:1 pre-read 5 min before the 1:1 Member tasks + team_memory & activity Auto-generate summary

Pattern 1 — Automated reports. Every morning at 9, a scheduler fires a trigger, and the LLM queries collaboration tool task states, git commits, and dashboard metrics all at once and writes the daily report. It goes out to the portal or the team messenger. "Portal" here means the in-house portal web covered in §20.3. server.py (FastAPI) is always running, and the View_*.html pages Claude wrote run on top of it, so the report slots in naturally as one more portal page.

Pattern 2 — Decision → automatic task creation. When a decision card (proposal P####) is registered, the card's implementation and verification items become collaboration tool tasks as is. Assignee and due date get filled in automatically. What the meeting settled with a "so that's what we'll do" lands on the board as cards without passing through anyone's hands.

Pattern 3 — Task → decision card reverse reference. This is Pattern 2 in the opposite direction. When a collaboration tool task is completed, MCP updates the original decision card's execution_log. "Was this decision actually executed?" is tracked automatically. Decision and execution are tied together in both directions (the dotted line in the figure).

Pattern 4 — Progress analysis. At quarter's end, pull every task from that quarter at once and analyze the delay patterns. "What kinds of tasks keep getting pushed back?" becomes input for the retrospective.

Pattern 5 — 1:1 pre-reads. Five minutes before a 1:1 meeting, synthesize that member's collaboration tool tasks, team_memory slot, and recent activity into a pre-read summary. Instead of burning five minutes on "what were you working on again?", the 1:1 starts with the substance.

The five patterns share one thing. Value is recovered only where a pattern reduces repetitive work that would otherwise stay in human hands. A pattern added to look impressive only adds operational burden.


20.4.4 Phased Adoption — Don't Connect Everything at Once

Connecting all five MCP servers on day one is the most common failure. There is an order.

Stage What Rough Duration
1 Install one MCP server (ClickUp or JIRA) 1–2 days
2 Pilot one pattern (automated reports) About 1 week
3 Run all 5 patterns 1–2 months
4 Develop your own MCP server (only if a niche tool requires it) 1–3 months

The durations above are the author's estimate (unverified), based on Project A's mid-sized (10–50 person) team. They vary with team size and tool familiarity. What is clear: stages 1 through 3 are enough for most teams. Go to stage 4 (developing your own server) only when you must connect a niche in-house tool that has no MCP server on the market. Most of the benefit is recovered without ever going that far.

There is a reason to mention JIRA separately. Where ClickUp is the team's internal board, JIRA is often the tool you share with external organizations — publishers, outsourcing partners. Connect a JIRA MCP and you can automatically pull the outsourcing board's progress before each weekly meeting and shortlist the tasks suspected of slipping. The meeting starts at "decisions" rather than "status sharing." But the more externally shared the tool, the heavier the permission and leak pitfalls of the next section weigh.


20.4.5 Four Pitfalls — Why MCP Cuts Both Ways

MCP is the act of putting tools in the LLM's hands. If what's in the hand is a knife, someone can get cut.

Pitfall 1 — Permission incidents. If the LLM also holds write permission, unintended changes happen — one casual "clean this up" bulk-flips task statuses. The prescription is clear. Start MCP servers read-only. Open up write only where it is truly needed, as in Patterns 2 and 3, and even then behind a confirmation gate before execution. The LLM asking the human about the 'in review' classification in the earlier transcript is the same spirit — when in doubt, stop.

Pitfall 2 — Data leaks. Company data pulled in via MCP gets transmitted to an external LLM API. Balance values, unannounced content, and revenue metrics can flow out as is. The prescription: for sensitive data, use a self-hosted LLM, or substitute placeholders at the MCP server layer before anything leaves (the same IP protection principle that runs through this book).

Pitfall 3 — Dependency explosion. Chain five MCP servers into the critical path, and when one goes down, the entire morning report stops. The prescription: keep only one or two core servers as critical and split the rest off as auxiliary. When an auxiliary server dies, only its section comes up empty — the report itself still ships.

Pitfall 4 — API cost explosion. A single MCP call burns LLM tokens and an external API call at the same time. Run the automated report every five minutes and the cost quietly balloons. The prescription: put a cap on call frequency, and cache query results that rarely change.

Tie the four pitfalls into one line, and MCP's safe position is: start read-only, keep only the core critical, filter sensitive data, and cap the calls.


20.4.6 Results — Where the Time Is Recovered

These are the changes I felt on Project A before and after running MCP. The time figures below are the author's empirical estimates (unverified) — read them as direction and ratio, not absolute values.

Item Without MCP With MCP Direction
Compiling info for the daily report Manual, 30–60 min Automated, \~5 min Large reduction
Syncing info across tools Manual human work Automatic Manual work eliminated
1:1 prep 10–15 min Auto summary, \~3 min Reduced
Tracking outsourcing progress Meetings only Real-time lookup Always available
Decision ↔ task linkage Manual Two-way automatic No more omissions

The biggest recovery comes from "time spent gathering information." Decisions and judgment are still human work. What MCP cuts is the stage before that — the grunt work of digging through scattered tools and pulling everything into one place. The three seconds at the top of this chapter are exactly that spot.


20.4.7 Closing Part 20

Part 20 built the team collaboration system in four layers.

Chapter Core
20.1 atom operations — category classification, quarterly cleanup
20.2 Team member memory — team_memory 5-person slots (leeminsoo, teammates A/b/c, shared), shared vs. personal separation
20.3 Portal web — always on via server.py (FastAPI), build_index.py, nginx, and nssm; runs View_*.html
20.4 MCP — 5 patterns, phased adoption, permissions first

The four chapters are not standalone tools; they are one body. The portal (20.3) is where the reports land; atoms and memory (20.1, 20.2) are where the LLM remembers team definitions like 'in review = incomplete'; and MCP (20.4) is the wiring that connects all of it to the collaboration tool and the documents. Detach any one of them and the value of the rest drops by half.

The next part covers governance and operations. The questions of safety, cost, copyright, and ethics that follow once you have connected this many tools — we take this chapter's pitfalls and elevate them into team-level rules.


Key Takeaways


20.4.8 Try It Yourself — Your First ClickUp MCP Connection

setup 1. Install a ClickUp MCP server (official or community) and issue a ClickUp API token. 2. Register the MCP server in your Claude Code settings, but start with a read-only scope only. 3. Check your workspace ID and limit the query scope to your own board.

prompt

From ClickUp, list the incomplete P0 tasks due this week, ordered by closest due date.
Add the assignee and the days left until the deadline, and put anything past due at the top.

verify 1. Open one of the returned tasks directly in ClickUp and cross-check the due date and assignee. 2. See whether the LLM asked back about ambiguous items on its own — if it asserted without asking, state the classification criteria explicitly and run it again. 3. Enter frequently used definitions ('in review = incomplete', etc.) as atoms in team_memory's shared slot so they apply automatically in future sessions.

Solo Scale-Down

If you are a solo developer with no team and no collaboration tool, MCP's first targets are GitHub and your local documents. Connect a GitHub MCP read-only and start with queries like "issues still open this week, oldest first," and instead of decision cards, query the markdown in your local decisions/ folder through a document MCP. Automated reports (Pattern 1) apply as is — just change the destination from a team messenger to your own notes file. The permission and cost pitfalls apply just the same to a team of one, so keep the read-only start and the call cap from day one.

Part 21 · Self Improving

Part 21 · Chapter 1. The Retrospective as the Starting Point of Everything

Friday, 6:40 p.m. I was about to close my laptop and head home when the day's work felt strangely familiar. I had tracked down and fixed a broken enum reference in a data sheet — and I was certain I had fixed the same thing last week. And the week before that. Each time I retyped the same prompt; each time I checked the same items in Claude's output. I only realized it was the third time just before the laptop lid closed.

That sensation — "this feels strangely familiar" — is the single most important moment in this entire book. Let it slip, and next week you repeat the same task a fourth time. Catch it and write it down as one line, and that line becomes a skill next week, and that skill becomes an atom a month later, injected automatically. The place where you catch it is the retrospective.

The other parts of this book covered "here is a tool" and "here is a pattern." This chapter covers where all those tools and patterns ignite. Where does a new slash command get made, how does a new atom get codified, and who weeds out the tool that gets used less than once a month? The answer always converges on the same place: the retrospective.


1.1 Catching Familiarity in One Line — A Worked Transcript

A retrospective is not a grand ritual. Let me open one of my actual daily retro files as is. The real path is retro/daily/YYYY-MM-DD.md; after a month, these files get compressed into weeklies (retro/weekly/YYYY-Www.md) and then into monthlies (retro/YYYY-MM.md). The rules live in retro/README.md.

The flow I actually use is not writing the retro file by hand. I hand the AI that day's work log and have it extract the patterns. Below is a reconstruction of an actual exchange from one Friday evening. The full prompt comes first, then Claude's raw output, unsummarized — including the parts that miss.

[Prompt — typed by the human]

Summarize today's work as a retrospective. What I did today: ① fixed a broken enum reference in the character data sheet (the grade column didn't match the enum sheet) ② unified the tone of NPC dialogue ③ reran relation-map. If any of these is a task I've repeated two or more times this week, flag it, and if there is a candidate worth turning into a tool, list it separately. If there's nothing, say so. Don't force it.

[Claude's raw output — no summary, no embellishment]

Here is today's work retrospective.

Repeated patterns detected - ① Fixing the broken enum reference → the same item appears in this week's Tuesday retro. Flagging it as repeated 2+ times. - ② Unifying NPC dialogue tone → first appearance this week. Not a repeat. - ③ Rerunning relation-map → you are already using the relation-map-gen skill, so this task is already tooled. Not an additional candidate.

Tool candidates - Enum reference validation looks like it could become a skill — something that cross-checks the data sheet's column values against the enum sheet and reports mismatches. - That said, I don't know the exact structure of the two sheets, so I can't be confident this is truly a repeatable pattern. You need to confirm whether Tuesday's breakage and today's breakage are the same kind of breakage. If the same column broke for the same reason, this is a strong skill candidate; if a different column broke for a different reason each time, it's hard to bundle into a skill.

The last paragraph is the key. The AI did not declare "let's build a skill." It asked back: a human should confirm whether it's the same kind of breakage. This is why the retrospective is not automatic.

[Human verification — taking the AI's suggestion, but filtering it]

I opened Tuesday's retro and compared. Tuesday's breakage was the grade column; today's breakage was also the grade column. The same kind. The AI's candidate suggestion passed verification. So I left one line in the daily retro.

Repeated task: enum-grade reference validation (Tue & Fri, 2 times) → skill candidate. Decide on promotion at the next weekly retro.

That one line is everything. It took less than five minutes. And that one line is the first beat of the self-improving loop. If item ③ — the one the AI flagged as already tooled — had also gone onto the candidate list, one more unused duplicate tool would have been floating around a month later. With both the AI's filter and the human's filter working, exactly one real candidate remained.


1.2 Why the Retrospective Is the Starting Point

Which tasks repeat, which tools get used often, which atoms are missing — none of this is visible from a single work session. In the transcript above, the enum breakage surfaced as a candidate not because of "today" but because "Tuesday and today" were overlaid. Patterns emerge only when you overlay traces accumulated over a week, a month, a quarter. The retrospective is the time you deliberately create that overlay.

Patterns found in a retrospective split into two branches.

You cannot make this call in the middle of work, because it breaks the work flow. In the moment of fixing the enum breakage, there is no room to ask "is this the third time?" The retrospective time you set aside separately is the seat of that judgment.

Whether a tool actually delivers value after it is built is measured in the same seat. A tool used once a month and a tool that saves an hour are not worth the same. Both the measurement and the retirement decision happen in the retrospective. Without it, tools only accumulate and never get cleaned up. A few years on, dozens of unused tools are getting in the way of search and operations.

To use a drawer analogy: the retrospective is the time you periodically empty your desk drawer. If the pen you use every day and the notepad you haven't pulled out in a year share one compartment, finding the pen costs a few extra seconds every time. Tools are exactly the same.


1.3 The Compression Flow of Daily, Weekly, and Monthly Retros

My retrospectives run on three tiers. The daily collects pattern seeds, the weekly bundles those seeds and compresses them into tool candidates, and the monthly evaluates each tool's economics and either codifies it as an asset or retires it. Each tier takes the output of the tier below as its input.

flowchart TD
    W["Do the work
(data sheets · design docs · automation)"] -->|material accumulates| D["Daily retro
retro/daily/*.md
5–10 min · 1–3 pattern seeds"] D -->|compress 5 dailies| WK["Weekly retro
retro/weekly/*.md
30–60 min · tool-candidate calls"] WK -->|synthesize 4 weeklies| M["Monthly retro
retro/YYYY-MM.md
1.5–2 hr · economics · retirement"] M -->|promote verified patterns| A["Permanent assets
feedback.md / workflows.md
+ atom registration"] A -->|auto-injected via JIT hook| W style D fill:#e3f2fd,stroke:#1565c0 style WK fill:#e8f5e9,stroke:#2e7d32 style M fill:#fff3e0,stroke:#ef6c00 style A fill:#f3e5f5,stroke:#6a1b9a

The last arrow closes the loop. Patterns codified as permanent assets are injected automatically into the next work session through a JIT (Just-In-Time) hook. In my environment, a hook called inject_memory.py picks out the relevant atoms and inserts them every time user input comes in. Once enum-grade validation is codified as an atom, the next time I type something like "validate the data sheet," that atom comes along on its own. The human no longer has to remember "right, there was that validation rule" every single time.

Take the retrospective out, and only the top-to-bottom arrows remain; the final arrow that returns assets to the work is severed. The loop does not close. The meaning of the word self-improving is precisely that this loop is turning. Tools improve the tools themselves; atoms grow more atoms. The engine is the one hour of retrospective a human sets aside.


1.4 The Five Things a Retrospective Sparks

From my impression after about half a year of running retrospectives on an MMORPG project I operate, a single retrospective sparks the following five kinds of output. The frequencies below are not precise statistics but my operational feel (author's estimate, unverified), and not every retrospective produces all five. On a quarterly scale, all five show up at least once.

Spelled out, the five are these. If you repeated the same decision two or more times this week, that is a new atom candidate. If you retyped the same prompt pattern several times in one week, that is a new skill candidate (the enum-grade validation in the previous section was this case). If a skill you used this week gave lackluster results, that is an existing-skill improvement — prompt adjustment, added verification, input standardization. An atom built last quarter that matched zero times over a month is a retirement candidate; unused, it only takes up tokens. Finally, economics evaluation weighs each tool's usage frequency against the manual effort it saves and decides keep, improve, or retire.

Here is how the five sit on one screen as a matrix. The horizontal axis is "does it repeat"; the vertical axis is "is it valuable."

Repetition frequency → high Outcome value → high High value · low repetition → Leave as is (hold off on tooling) High value · high repetition → New skill / new atom candidate (enum-grade validation goes here) Low value · low repetition → Ignore Low value · high repetition → Retire / simplify candidate

What a retrospective does, in the end, is scatter the week's tasks across these quadrants. What lands in the upper right becomes a tool; what lands in the lower right gets weeded out. This sorting is how self-improving actually operates.


1.5 Where Patterns Get Codified as atoms

Let me follow the enum-grade validation candidate from the previous section the rest of the way: promoted to a skill, and from there codified as an atom. An atom is the form a pattern takes when something roughly discovered in a retrospective has passed verification and become a permanent asset.

My memory already holds atoms codified that way. One of them is retro_atom_natural_invitation. As the name says, it carries the principle that "in a retrospective, an atom appears as a natural invitation, not a command." This atom is itself a meta-pattern discovered over many cycles of retrospectives — only after experiencing several times that compulsively forcing "this must be pinned down as an atom" mid-retro turns the retrospective into a formality did it harden into one line.

Whether codification actually pays off is also managed by score. My environment has a script called atom_score.py that grades how often each atom actually matches and gets used. The results are saved to _scores_latest.json, and atoms that clear a certain score are injected automatically into CLAUDE.md. In other words, the better used an atom is, the more often it surfaces in front of you, while an unused atom loses points and drifts toward the retirement pile. This grade-and-inject cycle is the automated portion of the quadrants in §1.4.

One thing needs to be said honestly here. This score does not convert directly into a quantitative metric like "saved 30 hours a month." The time an atom saves is tricky to measure. So rather than asserting an ROI (return on investment) figure, it is more honest to speak only in direction and proportion: "well-used atoms score high, and high-scoring atoms get injected more often and cut manual effort" — direction, not multiples.


1.6 What Happens Where There Is No Retrospective

The common scenery on a team that does not set retrospectives aside looks like this.

One hour of retrospective makes most of this scenery disappear. Think back to the Friday evening in §1.1 — the only difference is whether you write "this feels strangely familiar" down as one line or let it slip away. Writing it takes five minutes; the time lost by not writing it compounds into the fourth and fifth repetitions. It is the classic case of losing more time by refusing to spend any.

Of course, there is no need to stand up a grand retrospective system from day one. On a large team, starting from a five-minute daily retro may feel frustratingly slow. But laying down the full daily-weekly-monthly three-tier system in one go makes it easy to chase the form and miss the substance. The safe order is for someone to experience the value of retrospectives firsthand at the smallest tier, then pull it up to the next.

Every tool, atom, and pattern you meet in this book ignited in someone's retrospective. It is no exaggeration to say that this book itself is the product of half a year of my accumulated retrospectives. "The retrospective is the starting point" is not a metaphor — it points to the fact that this book's table of contents itself came out of retrospectives.


Beyond Games. A retrospective that catches the sensation of "this feels strangely familiar — I did this last week too" in one line is the entrance to self-improvement in any workplace with recurring tasks, not just game development. When you close out your day, write just one line — "what did I do by hand twice today" — and on Friday overlay the week's five lines; the pattern that each day kept hidden shows itself at the weekly scale. For example, if a general-affairs staffer catches "I retype the same form email every week" in a retrospective, that one line becomes an email template next week and an automated send rule a month later. The point is not the sophistication of the tool but the act of overlaying traces itself — and putting a filter on the AI when you have it extract candidates: "don't force it, and drop anything already automated."

1.7 Try It Yourself

This is the smallest version of adopting the retrospective loop for the first time. You can start with almost no tooling installed.

setup

  1. Create one folder, retro/daily/, inside your working folder.
  2. Open an empty file named with today's date, retro/daily/2026-06-06.md. Nothing more is needed.

prompt

At the end of each workday, give the AI that day's work log and ask like this.

Today I did [tasks 1, 2, 3]. If any of these is a task I've repeated two or more times this week, flag it, and if there is a repeating pattern worth turning into a tool (skill), list it separately. Drop any task that is already tooled from the candidates. If there's nothing, say so. Don't force it.

The last two sentences ("drop anything already tooled," "don't force it") are the filter. Without them, the AI over-generates plausible candidates every time, and a month later unused tools have piled up.

verify

  1. Check yourself, with one direct comparison, whether the repeated item the AI flagged really is the same kind of repetition (like the Tue/Fri grade column comparison earlier). Same kind: confirm the candidate. Different: discard.
  2. Leave the confirmed candidate as one line in the daily retro: Repeat: [task] (N times) → skill candidate, decide at weekly retro.
  3. A week later, hand the five daily retros back to the AI and ask it to "shortlist the candidates worth promoting to skills." Build tools only from candidates that have survived twice or more.

Solo Scale-Down

If you are starting alone, with no team and no extra tools, shrink it down like this.

The point is not the sophistication of the tool but the act of overlaying. A single day hides the pattern; a week reveals it. The five minutes you spend catching that reveal is the entrance to the self-improving loop.


Key Takeaways

Part 21 · Chapter 2. The Retrospective System and Atom Promotion — Turning Discoveries into Permanent Assets

Monday morning. I had pulled up last week's five daily retrospectives on one screen, about to start the week. Tuesday's entry said, "Forgot the integrity check before exporting the data sheets." Thursday's entry had almost the same sentence. And that very Monday morning, I was doing it again. I had pushed a sheet with broken FKs straight into the client/server build and had to pull it back out. It was the third time.

This moment is the heart of the retrospective system. The fact that you are doing something for the third time is never visible while you are doing it. Your hands move on habit and your head whispers, "this is just something I always do." Repetition becomes visible only when you gather the traces and look at them afterward. The retrospective is the device that gathers those traces, and atom promotion is the device that pins down the repetition you find there as a rule, so you never do it by hand again.

This chapter follows how those two devices interlock — how one actual daily retrospective file turns into one atom line in the JIT manifest, all the way to the end.


21.2.1 Discoveries Come Only from a Pile of Traces

There is a premise to establish first: repetition is not perceived in real time.

A game designer's day is a string of decisions. Which enum a data sheet column should use, whether a skill cooldown should be in seconds or in frames, where in the document to record a fuzzy agreement from a meeting. Each of these decisions is too small to stick in memory. But if I am making the same decision three times in one week, it is no longer a decision — it is a rule. And if it is a rule that I keep re-deciding from scratch every time, that is waste.

The problem is that this waste is invisible. So I leave traces. Five minutes a day: one line each for what I did today and for anything I did more than once today. After a week, five pages of traces have piled up, and only then does "huh, this is written down three times" become visible.

This is what separates a retrospective from a plain diary. A diary records impressions; a retrospective records traces in order to extract patterns. That is why the format must be fixed. If the format changes every time, you cannot lay five pages side by side and compare them, and if you cannot compare, you cannot see patterns.


21.2.2 The Three Cycles Are Units of Role, Not Units of Time

The reason retrospectives are split into daily, weekly, and monthly cycles is not the passage of time. It is that each cycle does fundamentally different work.

Daily · 5–10 min Role: pin down traces What I did today Decisions made twice Tools left unused Handoff to next session Output: 5 pages/week Weekly · 30–60 min Role: extract patterns Review 5 daily pages 3+ repeats → candidate Pin down pending- atom Output: 1–3 candidates Monthly · 1.5–2h Role: economics · promotion Tool economics review Promote / retire decisions Quarterly planning Output: official atoms

The daily pins down traces. No judgment — just write. The weekly bundles the five pages of traces and looks for patterns. This is where the first judgment — "this is a repetition" — comes in. The monthly looks at the entire accumulated toolset and evaluates its economics. It decides what to keep alive and what to throw away.

Drop one cycle and the others collapse. Weekly without daily means you cannot remember what happened a week ago, so the traces come up empty. Monthly without weekly means facing a month's worth of dailies in one sitting — comparing 22 pages at once is close to impossible. No patterns appear; only fatigue accumulates.

The workshop analogy fits well. The daily is the five minutes of clearing your desk every evening. The weekly is the thirty minutes of reorganizing one drawer on the weekend. The monthly is the two hours each quarter spent reviewing the layout of the whole workshop. If you never clear the desk, you cannot reorganize the drawer on the weekend, and if the drawers are a mess, staring at the layout gets you nowhere.


21.2.3 The Daily Retrospective — Pinning Down in Five Minutes

The daily retrospective files I actually use pile up by date at paths like retro/daily/2026-05-30.md. The /retro slash command lays down the template automatically.

# Daily Retrospective 2026-05-30

## What I Did Today (3–5 lines)
- Added 12 enum types to the new skill data sheet + reordered the cooldown columns
- Balance sim pass 1 (adjusted drop table weights)
- Refreshed the client/server data export builds together

## Repetition Spotted (if any)
- Forgot the integrity check before the data export build again → built with broken FKs → third time
- Ran the balance sim without pinning the seed — not reproducible (second time)

## Retirement Candidates
- Tools not used even once today: (recorded only for monthly accumulation)

## Handoff to the Next Session
- Fill the two broken FKs (skill→effect references) first, then rebuild
- Candidate: consider making the sim seed-pinning option the default

Five minutes is enough to fill it in. Because the format is fixed, I never have to wonder anew what to write. The slots are set; I only fill the slots.

The decisive slot here is "Repetition Spotted." It is allowed to stay empty. Most days it is empty. But when the awareness hits that I did the same thing twice today, I write one line. The example above — "Forgot the integrity check before the data export build again → third time" — is exactly that. That one line gets bundled into a pattern at the weekly retrospective a few days later, and a few weeks after that it gets pinned down as an atom or a skill.

Automatic capture reduces the manual work. When the git commit log, the atom change history, and the skill usage log are merged into the daily retrospective automatically, half of the "What I Did Today" slot is already filled. The human only adds what the git log cannot see — the awareness that "I did this again."

The last slot, "Handoff to the Next Session," is a note to tomorrow's me. With it, loading context at the start of a new session takes one or two minutes. Without it, I spend longer groping for "what was I in the middle of yesterday?" My own MEMORY.md actually maintains a separate "check first next session" section, which is the accumulated, higher-level version of this daily handoff.


21.2.4 The Weekly Retrospective — Where Patterns First Show Themselves

The weekly retrospective starts by putting the five daily pages on one screen. The files pile up at paths like retro/weekly/2026-W21.md.

# Weekly Retrospective 2026-W22 (5/25–5/31)

## Summary of This Week's Work
- Updated skill/balance data sheets, ran the drop table sim twice
- 4 data export builds (2 of them built with broken FKs/enums)

## Patterns Spotted
- "Forgot the pre-export integrity check" repeated in 3 dailies → atom candidate
- "Balance sim seed not pinned" repeated in 2 dailies → review sim defaults

## Atom Candidates
- pending-data-check-before-export (a rule that enforces integrity verification before export builds)

## Skill Candidates
- (None — an atom is enough this week)

## Existing Tool Check
- Unused: relation-map-gen (0 uses this week)
- Most used: check (integrity cascade), excel-reader, /retro

## Next Week's Plan
- Run pending-data-check-before-export one more week, then decide on promotion

This is where judgment enters for the first time. "Forgot the pre-export check, repeated in 3 dailies" is an arithmetic fact, but "this is worth pinning down as an atom" is a judgment. The reason three repetitions is the baseline is simple. Once is chance, twice might be chance, three times is a pattern.

Once the judgment is made, I pin it down immediately — not as an official atom, but as a provisional atom carrying the pending- prefix. It lands in my project memory folder like this.

~/.claude/projects/<project>/memory/
  pending-data-check-before-export.md

The pending- prefix is a marker that says "this is still under verification." The marker matters because turning an unverified hunch straight into a team-wide rule breaks two things. One is trust — when unverified rules keep being wrong, people stop trusting the rules themselves. The other is accumulation — without a verification gate, hunches pile up as is and the memory becomes a garbage can.

So a pending- atom is run inside real work for a week, at most a month. If it genuinely proves useful every time, it survives; if it never once applies, it is quietly deleted.


21.2.5 The Monthly Retrospective — Measuring Tool Health and Choosing What to Keep

The monthly retrospective is where I spread out a month's accumulation and check the health of the entire toolset. The files pile up by month, like retro/2026-05.md.

# Monthly Retrospective 2026-05

## This Month's Totals
- Daily retrospectives: 22, weekly retrospectives: 4
- New atoms: 4 (data-check-before-export, sim-seed-pinning, and others)
- New skills: 1 (relation-map-gen option upgrade)
- Retired atoms: 1

## Tool Economics Review
- Monthly use count per skill + felt savings (qualitative)
- Skills used less than once a month → retirement candidates
- Most valuable tools: check (integrity cascade), excel-reader, /retro

## Atom Distribution
- Totals by prefix (data: X, sim: Y, meeting: Z ...)
- Retirement candidates: atoms with 0 matches in a month

## Quarterly Plan
- To introduce next month: impact (impact tracking), automatic schema-doc refresh

## Book Material (when applicable)
- Cases this month worth citing in the book: 1 worked example of atom promotion

The heart of the monthly is the economics review. Every tool looks valuable when you build it, but a month later half of them go untouched. I use five yardsticks to sort them out.

The five criteria are usage frequency, time saved, cognitive load, maintenance cost, and replaceability. Usage frequency: once a month or more and it stays for now; less than that and it goes on the retirement list. Time saved: multiply the felt savings per use by the frequency — I do not commit to minute-level numbers here. "It feels like a few minutes per use, and I use it ten times a month, so the total is large" is the honest level of qualitative judgment. Cognitive load: when the slash commands I have to memorize exceed twelve, I treat it as a signal to clean up. There is a limit to how many commands a person can carry around in their head. Maintenance cost: is this a tool I have to touch every time the data sheets change? Replaceability: has a simpler method appeared?

Combining the five yardsticks, I decide whether to keep, drop, or fix each tool. An hour is enough to go through thirty to fifty tools. That said, in the first month or two there are so few tools that the review takes well under an hour. In that period it is better to spend the time creating new tools than evaluating them. The economics review is a safety device that only starts doing real work once tools have accumulated.


21.2.6 Worked Transcript: From One Daily Line to One Atom Line in the JIT Manifest

From here on is the spine of this chapter. I follow the process in which the pieces scattered above actually start from a single line and travel all the way to automatic injection in every session — from the retrospective moment to the very end. The moments where I rejected the output and where the AI answered ambiguously are left in as they happened.

Step 1 — Awareness in the daily. While writing the May 30 daily retrospective, I put one line in the "Repetition Spotted" slot.

- Forgot the integrity check before the data export build again → built with broken FKs → third time

I built no tool at this point. I only left a trace.

Step 2 — Confirm the pattern in the weekly and ask the AI to turn it into an atom. A few days later, when I spread the five daily pages at the weekly retrospective, the same line was there in three places. So the full prompt I threw at Claude was this.

Three of this week's daily retrospectives repeat "I keep forgetting to run the integrity check by hand before the data export build." I want to turn this into a pending atom. Draft the atom body for me. Format it like the other atoms in our memory: frontmatter (name, status, trigger) plus a one-paragraph rule.

Claude's raw output was this (verbatim, no summarizing).

---
name: data-check-before-export
status: pending
trigger: ["data", "export", "pre-build", "integrity"]
---
Before running a client/server export build of the game data sheets, always
run the integrity check (check cascade) first, and export only after
confirming zero FK, enum, and reference violations. If a violation is found,
stop the export and notify the owner of the failure details via ClickUp.

You can save this atom as pending-data-check-before-export.md. One caveat: among the trigger keywords, "data" is too broad and could also match situations unrelated to export (browsing data sheets, data meetings, and so on). I recommend removing it.

Step 3 — Human verification and partial rejection. The AI doubting one of its own triggers was a fair point. But I went one step further. I did not like that the rule body bundled both "integrity verification" and "ClickUp notification on failure." The failure notification is already a separate atom owned by integrity_check_clickup_notify. If one atom carries two responsibilities, then later, when I want to change only the notification rule, I end up touching the verification rule too. So I rejected it and asked again.

Good — let's drop the "data" trigger. And the ClickUp notification on failure is already handled by integrity_check_clickup_notify, so take it out of this atom. Narrow this atom down so it is responsible only for pre-export integrity verification.

After the re-request, the body Claude narrowed down was this.

---
name: data-check-before-export
status: pending
trigger: ["data build", "export", "pre-build verification"]
---
Before running a client/server export build of the game data sheets, run the
integrity check (check cascade) first. Checks: FK referential integrity, enum
value validity, cross-table reference consistency. Export only after zero
violations are confirmed. Failure notification is the responsibility of integrity_check_clickup_notify; this rule does not get involved.

This is the version I accepted. It carries a single responsibility, and the triggers are narrowed to the export context. I saved this file as pending-data-check-before-export.md. The key point here is that I did not use the AI's first output as is. The AI supplied the format and doubted one trigger, but the design judgment — "narrow it to one responsibility" — was made by a human.

Step 4 — One week of operational verification. Over the next week, this pending atom came up every time I ran a data export build, and it actually caught broken enums right before a build twice. It never once misfired. It had earned the right to survive.

Step 5 — Promotion decision and score measurement at the monthly. At the monthly retrospective, I put this pending atom on the promotion list. Promotion is judged by measurement, not gut feeling. My environment has a script that scores each atom's match frequency and usefulness.

python ~/.claude/scripts/atom_score.py
# → refreshes ~/.claude/projects/<project>/memory/_scores_latest.json

The script tallies how many times each atom matched its triggers over the period and whether it was actually cited in the work, and drops the result into _scores_latest.json. Atoms whose score crosses a set threshold are wired to be auto-injected into CLAUDE.md. By actual measurement in May 2026, pending-data-check-before-export matched on every export build for a week, so the score was sufficient. Promotion confirmed.

Step 6 — Remove pending-, register in the JIT manifest. I stripped the prefix, made it an official atom, and added one line to the JIT manifest.

~/.claude/projects/<project>/memory/_jit_manifest.json

The UserPromptSubmit hook (~/.claude/hooks/inject_memory.py) reads this manifest on every input. If the input contains "data build" or "export", the atom body is automatically inserted into the context.

Step 7 — The loop closes. The moment I typed "run the data export build" for the next build, without my asking for anything, Claude spoke first.

Shall I run the integrity check (check cascade) before the export? I will check FK references, enum values, and cross-table reference consistency, confirm zero violations, and then export.

The one line of "I did this again," written in a daily retrospective three weeks earlier, had come back as a rule that guards the current work on its own. The verification I used to do by hand, I never do by hand again. This scene is exactly what "the loop closed" means.


21.2.7 From Discovery to Auto-Injection: The Full Picture of the Promotion Loop

Compressed into a single flowchart, the worked transcript above looks like this. Discovery happens in the daily, verification is done by the operating period, promotion is decided by measurement, and the manifest finishes turning it into an asset.

flowchart TD
    A["Daily retrospective
One line: 'I did this again'"] --> B["Weekly retrospective
Review 5 pages for patterns (3+ repeats)"] B --> C["Ask AI to draft the atom
→ raw output → human review/rejection → re-request"] C --> D["Pin down pending- atom
~/.claude/.../memory/pending-*.md"] D --> E["1–4 weeks of verification in real work"] E -->|Never once applied| X["Quietly retired"] E -->|Useful every time| F["atom_score.py measurement
_scores_latest.json"] F --> G["Promotion decision at the monthly retro
remove pending-"] G --> H["Register in JIT manifest
_jit_manifest.json"] H --> I["UserPromptSubmit hook
inject_memory.py auto-injection"] I --> J["Next session: on keyword input,
a past discovery guards the current work"] J -.Discovered again in retro.-> A classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; classDef fail fill:#fee2e2,stroke:#dc2626,color:#7f1d1d; class F,I code; class A,B,G human; class D,H data; class J pass; class X fail;

The final dotted line is the whole point of this diagram. An auto-injected atom exposes yet another repetition, which goes back into the retrospective and gives birth to the next atom. Each turn of the loop removes one more thing done by hand. Let this loop accumulate for six months to a year, and the retrospective is no longer a diary — it is the brain of the work system.


21.2.8 Let the Retrospective Invite Atoms Naturally

The most fragile link in the promotion loop is step 1 — the moment of writing "I did this again." When busy, people leave the retrospective slot empty and move on. Then no trace remains; without traces, no pattern shows up at the weekly; without patterns, no atom is born. The entrance to the loop gets blocked.

That is why my environment has an atom called retro_atom_natural_invitation. The rule: when writing a retrospective, do not force atom creation as a duty — leave it as a natural invitation. That is, not "you must extract one atom candidate today," but: keep the "Repetition Spotted" slot in the template as a slot that may stay empty, while gently nudging you to drop in a line when there is something worth one line. Make it a duty and you end up squeezing out fake patterns; leave it as an invitation and only real repetitions get caught, naturally.

This hair's-breadth difference decides whether the loop is sustainable. A mandatory retrospective does not survive two weeks before it fills up with perfunctory lies. An invitation-style retrospective leaves the slot blank on days with nothing to write, so it carries no burden, and therefore it lasts. It has to last for traces to pile up, and traces have to pile up for patterns to show.

This atom itself was born from a retrospective. While running retrospectives as a duty, I noticed within days — in the dailies — that the slots were being filled with fakes; that discovery went through the weekly and was promoted into this rule. A rule that improves the retrospective came out of the retrospective.


21.2.9 Where Things Usually Break

Run the loop for a while and it keeps collapsing in the same places.

Skipping retrospectives is the most common. Skip three days because you are busy, and those three days of traces are gone forever. The countermeasure is simple — leave every other slot empty if you must, but write the one line of "What I Did Today." Even one minute instead of five still leaves a trace.

Re-inventing the format every time is also dangerous. Free-form entries cannot be laid side by side and compared across five pages. If comparison fails, the weekly's core job — pattern extraction — becomes impossible. That is why /retro forcibly lays down the template.

Not retiring tools is another trap. If atoms and skills only grow and never get dropped, cognitive load accumulates. The moment slash commands exceed twelve, your head can no longer hold all the tools. The monthly economics review is the only device that stops this accumulation.

Skipping the promotion gate is dangerous too. Turn hunches straight into official atoms and unverified rules pile up. The gate — passing through pending- and promoting by measurement — must sit between discovery and asset-making.

Finally, skipping retrospectives because you work alone and have no team to share with is a misunderstanding. The entire worked transcript above ran in a solo environment. Only the team-share merge step drops out; the discovery → pending → measurement → promotion → JIT injection loop runs exactly the same solo. If anything, in a solo environment this loop is the only external reviewer you have.


Beyond Games. The promotion loop — one daily line passing through verification to become a permanent rule — is a procedure for making sure a lesson learned once never has to be done by hand again, whatever your line of work. The key is the gate: instead of fixing a discovery into a team rule right away, run it as pending for a week and formalize it only when it actually proved useful every time — because turning unverified hunches straight into rules makes people stop trusting the rules themselves. For example, if an operations team keeps "double-check that the totals add up before submitting the report" as a provisional checklist, runs it for a week, and elevates it to an official standard procedure only after it actually catches an error or two, the design judgment of narrowing each rule to a single responsibility (not bundling several checks into one line) follows naturally.

Try It Yourself

setup. Lay down the retrospective folders and template.

mkdir -p ~/.claude/projects/<your-project>/memory/retro/daily
mkdir -p ~/.claude/projects/<your-project>/memory/retro/weekly
# Save one daily template file as retro/_template_daily.md

prompt. After a week of daily retrospectives has piled up, throw this at Claude in your weekly retrospective.

I'm pasting this week's five daily retrospectives. Find the tasks and decisions repeated three or more times and organize them as atom candidates. Each candidate should have frontmatter (name, status: pending, an array of trigger keywords) and a one-paragraph rule. If a trigger is too broad, propose a narrower one; if an atom carries two responsibilities, propose splitting it.

verify. Do not use the candidates as is — verify three things. (1) Does each atom carry exactly one responsibility? If two, reject and ask for a split. (2) Do the trigger keywords match only that task's context? If too broad, reject. (3) Was it really repeated three times, or just twice by chance? If chance, do not even create the pending file. Save only the candidates that pass all three checks with the pending- prefix, and promote them officially only after a week of operation in which they proved useful every time.


21.2.10 Solo Scale-Down

If you have no team, no JIT hook, and no score script yet, you can imitate the entire loop with the following single file.

Create one file, retro.md, with just three slots.

## Today (1 line)
- 

## Did It Again (1 line, if any)
- 

## Pin-Down Candidates (when "did it again" hits 3, move it here)
- [ ] (one-sentence rule) — verified: useful ___ times

Fill in only the top two slots every day. When the same line piles up three times under "Did It Again," move it to the third slot and write it as a one-sentence rule. Count how many times that rule actually proved useful over the next week, write the count in the slot, and if it reaches three or more, move the sentence officially into your project memory (CLAUDE.md). Even without a JIT hook, a rule placed in CLAUDE.md follows along in every next session, and that alone completes the minimal form of the loop where past discoveries help present work.

What matters is not how fancy the tools are but that the gate exists. With just the gate — "did it again ×3 → pin it down → useful ×3 → make it permanent" — the self-improving loop runs on a single file.


Key Takeaways

Next Chapter Preview

Part 21 · Chapter 3. Closing the Self-Improving Loop

Does the cycle that starts in a retrospective come back to the retrospective? If it doesn't close, it is a memo, not a system.


I open a retrospective from six months ago. "Terminology isn't consistent." "Documents are hard to find." "I keep getting the same questions." I open the retrospective I wrote this morning. "Terminology isn't consistent." "Documents are hard to find." "I keep getting the same questions."

Word for word, identical. It's not that we skipped retrospectives. We did them faithfully for six months straight. Notion pages piled up, and at the quarterly workshop sticky notes covered the whiteboard. Yet what gets written keeps circling in place. It's not that the retrospectives didn't work. The loop never closed.

This is the last chapter of the book, so it takes on the last question. All the tools built along the way — the city generator in Part 6, the mobile review atoms in Part 14, the cost standards in Part 22 — what more do they need to become systems that grow on their own, rather than disposable things built once and finished? There is one answer: a closed circuit in which a statement from a retrospective takes effect automatically from the next session on, and that effect is measured and fed back into the retrospective. The mechanism that closes this circuit is the self-improving loop.


21.3.1 The Anatomy of an Unclosed Loop

A retrospective produces a statement: "We have too many meetings." A good statement. But it ends up as one line on a Notion page. Next week there are still too many meetings, and the next retrospective records the same line again. Between the statement and the improvement sits human memory. People forget. So the chain breaks.

To deserve the name self-improving, a retrospective statement has to turn into automatic behavior in the next session without passing through human memory. Four conditions have to hold.

First, the statement must be converted into an immediately executable form — not an abstract resolution, but one of: a skill, an atom, a manifest entry, or a slash command. Second, it must fire automatically from the next session on, with no one having to remember it. Third, the next retrospective must measure what it actually changed. Fourth, that measurement must circulate back as input to the next improvement.

When all four connect automatically, the loop closes. Patch even one step with "I'll remember to apply it next week," and the loop reopens at that very spot. And the next retrospective records the same statement again.

In drawer terms, it looks like this. If the retrospective ends with a memo — "we never use this pen, let's take it out" — the pen is still there next week. The loop closes only when a hand actually removes it, not when the memo is written. And only a re-check next quarter keeps unused pens from piling up in that spot again. The memo is the statement, the hand reaching in is the auto-trigger, and next quarter's check is the measurement. Drop any one of the three and the drawer gets messy again.


21.3.2 The Closed Shape of the Loop

Drawn out, the whole flow forms a closed circuit. It starts at the retrospective and ends there too.

flowchart LR
    A["Retrospective
(daily / weekly / monthly)"] -->|statement| B["Candidate identification
skill / atom / command"] B -->|quantification| C["Economic evaluation
ROI formula"] C -->|pass| D["Implement & register
guarantee auto-trigger"] C -.->|fail| F["Discard / hold
move to another form"] D --> E["Next session
auto-trigger"] E -->|accumulated measurements| A F -.->|record| A style A fill:#2d4a3e,color:#fff style E fill:#3e2d4a,color:#fff

The arrows go full circle and come back into the retrospective. That closure is the point. Each step's output becomes the next step's input, and the final measurement becomes input to the next retrospective. Insert human memory anywhere in between and that arrow breaks, and the circuit breaks with it.

Note that the dotted arrow — candidates that fall short on ROI dropping into discard or hold — also leads back to the retrospective. The judgment "this wasn't worth building" itself becomes a record in the next retrospective, and grounds for filtering the same candidate quickly when it comes up again. Discarding lives inside the loop too.

Statements that lead from a retrospective into self-improvement follow five set patterns (covered in §21.1.4): a skill to build, a skill to improve, an atom to build, an atom to improve, and an economic re-evaluation. Build these five into the retrospective template as slots and no statement slips through.

## Retrospective (Daily) — 2026-06-06

### 1. Today's Work
- (work summary)

### 2. Self-Improving Statements (5 Slots)
- Skill to build: <"none" if empty>
- Skill to improve: <>
- atom to build: <>
- atom to improve: <>
- Economic re-evaluation: <>

### 3. To Measure in the Next Retrospective
- <>

Empty slots are fine. The emptiness itself is a record that says "no new improvements today." But if all five slots stay empty for days in a row, that's not a lack of things to improve — it's a sign the retrospective is hardening into a ritual. That's when I throw in the trigger question: "What did I do by hand twice this week?"

Statements come out vague. "Meeting notes are too long." To grow one into a candidate, quantify it as a single deliverable. "Meeting notes are too long" converts into a meeting_summary skill — one tool that takes meeting notes and extracts only the decisions and action items. "The terminology confuses me" converts into a glossary_lookup atom holding 30 domain terms; "I get the same questions every time" into an /onboarding slash command that automates a new hire's first-day orientation. "Sync keeps getting missed" comes down to a manifest update plus a new JIT atom.

A candidate moves to the next step only once it is defined as "one specific deliverable." "Let's improve things overall" is not a candidate. A statement that can't be converted into a single deliverable can't be placed on the ROI scale, and what can't be placed there stops right there.


21.3.3 ROI Is a Question of Orders of Magnitude

Having a candidate doesn't mean building it. Before building, measure the return on investment. The formula is simple.

ROI Formula Time saved × Trigger frequency × Operating period Build time + Maintenance burden ROI =

Each term has a unit and a passing bar. Time saved is the human time eliminated per trigger, counted in minutes. Trigger frequency is the estimated count per week — once a week or more keeps a candidate alive. Operating period is the expected number of weeks until retirement — a tool that won't survive four weeks has a weak case for being built. Build time is the time for the first implementation and verification; maintenance is the monthly time spent on checks and fixes.

The numerator is cumulative savings; the denominator is cumulative cost. The resulting value drives the decision.

ROI Value Decision
10 or higher Build immediately
3–10 Build within a week
1–3 Hold as pending, re-evaluate in a month
Below 1 Discard in this form; consider another approach

ROI below 1 doesn't mean "this idea is useless" — it means "don't build it in this form." First check whether a single, lighter atom line could replace it, or whether a wrapper that only changes the entry point of an existing tool could solve it. Demoting what would have been a heavyweight skill to a one-line atom often cuts the denominator to a tenth and brings the ROI back to life.

Let me plug in real numbers. Take the JIT atom injection system I built on my personal PC on May 23, 2026 — infrastructure where a UserPromptSubmit hook reads the user's input and automatically injects the relevant memory fragments (atoms) — and run its ROI.

Time saved:        about 3–5 min per session (no more finding and invoking the relevant atom by hand)
Trigger frequency: 15–25 sessions per week (on my personal PC)
Operating period:  1+ year expected (it's infrastructure, so retirement is unlikely)
Build time:        4 hours (hook + manifest + atom verification)
Maintenance:       0.5 hours per month (adding and editing atoms)

ROI = (4 min × 20 times/week × 52 weeks) / (4 hours × 60 min + 0.5 hours × 12 months × 60 min)
    = 4,160 min / (240 min + 360 min)
    = 4,160 min / 600 min
    ≈ 6.9  →  "build immediately" band. The decision is backed by the formula

Here is where I owe you some honesty. The numbers above — 3–5 minutes per session, 15–25 sessions per week — are estimates based on my operating experience, not precise measurements. No stopwatch was involved. So the 6.9 isn't a value to trust down to the decimal point either.

And that's fine, because the ROI formula is a tool for reading orders of magnitude, not precision. If the result lands around 7, build it. Around 0.3, think again. Telling those two apart needs no decimal places. What matters is that even the decision not to build comes from the formula, not from gut feeling. "Not building it — the order of magnitude doesn't work out": once that one line is in the retrospective, the same candidate coming up again costs no further deliberation.


21.3.4 Building It Is Not the End — Registration and Trigger Verification

When a candidate passes, build it. But building is only half the job. The other half is registering it so it fires automatically from the next session on. Skip the registration and the tool exists, but it sits where no one's hands can reach it — and the loop breaks right there.

Each deliverable type registers in a different place. A global skill goes into ~/.claude/skills/, together with a guide atom describing how to use it. A project skill goes into that project's .claude/skills/. A new atom goes into the right folder, gets one line added to the MEMORY.md index, and gets a trigger registered in the JIT manifest — all three, or auto-injection never comes alive. A slash command goes into ~/.claude/commands/; a wrapper changes the entry point of an existing tool and gets a guide atom attached.

Miss the registration and the next retrospective produces another statement: "I built this — why isn't it being used?" That's not a new improvement statement; it's a bug report. You are rediscovering, in the retrospective, the registration you yourself skipped.

Even with registration done, one step remains: open a new session and confirm the thing actually fires on the intended trigger.

1. Start a new session
2. Type the trigger (e.g., "How is my family's health doing")
3. Check the JIT log → was the intended atom actually injected?
   (~/.claude/hooks/_injection_log.txt)
4. If it didn't fire → broaden the trigger regex in the manifest
   or add a manual invocation path

Skip this verification and you get repeat incidents of "I thought it was there, but it didn't come up when I actually needed it." Registration and triggering are different things. Registration is placing the file; triggering is the trigger actually catching. If the trigger regex is off by a single character, the tool is registered but never surfaces.


21.3.5 Measurement — The Arrow Back to the Retrospective

Run the new tool for a week to a month, then measure. This measurement is the loop's final arrow — the one that goes back into the retrospective.

Actual trigger counts come from the JIT log or the command invocation log. Actual time saved gets recorded in the retrospective in the form "a task that used to take N minutes finished in M." Side effects — false triggers, needless context pollution — get checked alongside. Then put the initially estimated ROI and the measured ROI side by side.

If the estimated ROI was 6 and the measured ROI is 0.8, discard without mercy. The system's cleanliness outranks the builder's pride. Unused tools piling up in the manifest become noise that eats away at the accuracy of the next retrospective.

But check once before pressing the discard button. The trigger regex may have been too narrow for it to fire at all, or it may simply have been forgotten for lack of a manual invocation path. First separate the genuinely worthless tool from the good tool whose trigger path was blocked. Throw away the former; unblock the latter.

Discarding, too, is decided in the retrospective. The decision "this tool is retired" is itself a self-improving deliverable. A cycle that only builds and never clears out is a cycle of monotonic growth, and a monotonically growing system is eventually crushed under its own weight.


21.3.6 The Marks of a Closed Loop

Four signals tell you the loop is closed.

First, the same statement doesn't repeat. If an item written once in a retrospective shows up a second time, something failed the first time around — candidate identification or implementation. The "terminology isn't consistent" from the top of this chapter, repeating for six months — that was the clearest evidence of an open loop.

Second, the counts of manifest entries and atoms don't only go up. Retirement happens. A healthy cycle clears out around 10–20% per quarter. A system that has never shrunk is a drawer that has never been cleaned.

Third, retrospectives get shorter. When the system runs well, the time spent groping for "what did I do yesterday" disappears, and five minutes becomes enough to fill the five statement slots.

Fourth, a new hire can take part in the retrospective within a week. That's possible when the retrospective format is standardized and the atoms and skills are visible.

The loop always breaks in the same places. Collect the failure modes, and the next time the same symptom appears, you can pick up the prescription on the spot.

Break Point Symptom Prescription
No statements All 5 slots empty every time Add the trigger question: "What did I do by hand twice?"
Doesn't resolve into a candidate Vague "improve things overall" talk Force quantification into one deliverable
ROI evaluation skipped Build first, ask later Turn the ROI formula into a 5-minute template
Built but never surfaces Registration missed Enforce the registration checklist
Surfaces but goes unused Trigger missing or misconfigured Broaden the regex and provide a manual path at the same time
No measurement No measurement slot in the retrospective Add a "to measure in the next retrospective" slot

Each failure mode gets stated in a retrospective, and that statement becomes input to self-improvement again. Even fixing the loop happens inside the loop. It's a meta-loop.


21.3.7 The Last Sentence of This Book

This book has been long. It started with information architecture, built a tool that generates cities, designed combat systems, automated mobile review, standardized costs, and drew atoms up out of retrospectives. This final chapter is the question that all those chapters' tools gather in one place to answer: does what you built grow on its own?

Self-improving comes down, in the end, to a single sentence.

What was decided in a retrospective takes effect automatically from the next session on, and that effect is measured and returned to the retrospective.

Without automatic effect, a retrospective is a diary. A well-kept diary is consoling, but it doesn't change a system. With automatic effect, the retrospective becomes the system's brain. Each day's statements change each day's behavior, and the results of that behavior make the next statements more accurate.

Every field this book covered — information design, systems, combat, mobile, cost, and Layer decomposition as the precondition for procedural generation and automation — evolves on top of this self-improving loop. Tools age, models change, projects end. But as long as the loop stays closed, the system is a little better today than it was yesterday. That is the last thing this book leaves you: not how to build tools, but how to make tools grow on their own.

May your next retrospective be the first turn of that loop.


Key Takeaways


Beyond Games. If "terminology isn't consistent / documents are hard to find" has appeared in your retrospectives, word for word, for six months, it's not that you skipped retrospectives — the loop never closed, because human memory sits between the statement and the improvement. The conditions for a closed loop are the same in any department: the statement resolves into one immediately executable deliverable (a template, a checklist, an automation rule), it works from then on without anyone having to remember it, and its effect is measured and fed back. For example, the statement "meetings run too long" converts into "one tool that takes meeting notes and extracts only decisions and to-dos," and before building it you check only the order of magnitude — (time saved × trigger frequency × operating period) ÷ (build and maintenance time) — to decide whether to build now or hold. Even the decision not to build has to come from the formula rather than from intuition, so that when the same candidate comes up again, it costs no further deliberation.

Try It Yourself

setup

  1. Add the five self-improving slots and a "to measure in the next retrospective" slot to your retrospective template.
  2. Pin the one-line ROI formula and the decision band table (10 and up: immediate / 3–10: within a week / 1–3: hold / below 1: discard) at the top of your retrospective file.
  3. Prepare a registration checklist (where each type registers: skill, atom, command, wrapper).

prompt

Fill in the five self-improving slots for today's retrospective.
Quantify each statement as "one deliverable," and for each candidate estimate the ROI as
(minutes saved × triggers per week × operating weeks) / (build minutes + maintenance minutes),
then attach a decision band (immediate / within a week / hold / discard).
State the basis for each estimated number in one line, and mark it "estimate" if it is not a precise measurement.

verify

  1. After implementing a passing candidate, open a new session and type the intended trigger.
  2. Check the trigger log — did the intended atom or command actually come up?
  3. One week to one month later, compare the measured ROI against the estimate in your retrospective; if it's below 0.8, first rule out a blocked trigger path, then decide on retirement.

Solo Scale-Down

You don't need a team. Alone, shrink it to this. One memo line at the end of the day — "What did I do by hand twice today?" Turn that one thing into one line of automation the next day (an atom, an alias, a snippet). A week later, check only whether that line was actually used. If it was, keep it; if not, delete it. One line of statement → one line of automation → one line of measurement. Those three lines are the loop's smallest unit.

Part 22 · Governance

22.1 Prompt Engineering — The Game Designer's One-Page Work Order

Primary readers: game designers pulling LLMs into production work (mid-size teams of 10–50) Scaled-down version for solo/hobbyist readers: §22.1.7, "If You're Solo, Just This Much"

Hoping to get three lines of NPC dialogue, I once typed "write five lines for this NPC." What came back were five lines that would have sat comfortably in any fantasy game — and therefore fit nowhere in ours. The tone was empty, the lines had no idea who this NPC was, and they did not connect to the dialogue around them. Each line, taken alone, was grammatically fine. The problem: reviewing those five lines took longer than writing them myself from scratch.

This chapter is about turning that one-line instruction into a one-page work order. General prompt-writing advice fills plenty of other books. Here I show the four things a game designer needs in hand when sitting down in front of an LLM — context, output format, hallucination guard, verification request — not as abstract fragments but as one npc_dialogue prompt that actually ran. We follow one full cycle: what went into that prompt, what came out, and what got rejected.


22.1.1 A Prompt Is a Work Order — All Four Principles Fit on One Page

A good work order is not short. Hand a new hire a task with "do your best" and you get a different result every time; hand an LLM "write some dialogue" and you get the generic-RPG average every time. The same model splits on the instruction sheet — that output quality differs by multiples is industry common sense, and this book does not promise that multiplier as a number. The direction, though, is clear: a prompt loaded with context and constraints produces output that is cheaper to review than a bare one-liner.

The four things a game designer's prompt must satisfy at once are these.

Principle One-Line Definition If You Skip It
① Context Give it what to look at when answering (vision · voice · adjacent lines) You get the generic-fantasy average
② Output format Nail down count, length, labels, and what is banned Review sprawls into interpreting free-form prose
③ Hallucination guard State explicitly: "invent nothing beyond the given material" It fabricates lore that does not exist
④ Verification request Make the output mark for itself which criteria it meets No grounds to pass it through the gate

Memorize these as four separate rules and one or two keep dropping out. So this chapter's approach is to build the four principles into a single prompt page as slots. When a slot sits empty, the missing principle is visible. The next section shows that one page whole.


22.1.2 [Worked Transcript] One Page of the npc_dialogue Prompt

This is prompts/narrative/npc_dialogue_v3.txt, actually in operation on my project (a mobile-first MMORPG, "Project A" hereafter), anonymized and carried over as is. City and NPC names and company-specific terms were swapped for the book, and the output is a reconstruction of the actual session. The input prompt is in a form you can copy and use directly.

Step 1 — Context Input: Start with Who This NPC Is

First, fill the slots with the material the prompt will reference. None of the three is written fresh — all are pulled from existing assets.

# Slot input (attached above the prompt body)
L0_vision:        # Cached — not resent on every call
  world_premise:  "A confederation of scholars' city-states where the mana seals are cooling"
  tone_manifesto: "Restrain sentiment. Characters reveal emotion through actions and objects, not by explaining it."
voice_profile:    # This NPC's identity (5 items)
  id: npc_doren_vale
  age_range: "50s"
  speech_pattern: "Speaks only in numbers. Almost never uses adjectives."
  world_knowledge: "Has recorded the seal vein's micro-tremors for 30 years. Knows nothing of affairs outside the scholars' guild."
  taboos:  "No mysticism vocabulary — prophecy, fate, gods (the city's tone is scholarly_strict)"
  relationship:  "Treats the player as an 'external variable outside the observation'; little wariness, little goodwill"
adjacent_lines:        # Immediate context — lines already spoken in this scene
  - (Player) "The bell tower's light stayed on all night — what's going on?"

The five voice_profile items here are the core of Principle ①. Age, speech pattern, scope of knowledge, taboos, relationship — these five are what make "Doren Vale" distinguishable from every other NPC. The scope of world_knowledge in particular (knows nothing outside the guild) is the advance work for Principle ③, the hallucination guard. Only by stating what he does not know do you keep the AI from stepping outside it.

Step 2 — The Prompt Body: Pinning Down Format, Hallucination, and Verification on One Page

[L0 context] world_premise + tone_manifesto                    (cached)
[voice_profile] npc_doren_vale, 5 items (yaml above)
[adjacent lines] the player's previous question, 1 line

Looking at the material above, write doren_vale's reply to the player's question.

[Output format — Principle ②] Exactly 3 lines, one per line, each 40 characters or fewer. End each line with one emotion label from (detached|wary|cynical). No other commentary or quotation marks.

[Hallucination guard — Principle ③] Use only what is in the material above; do not invent new facts about the seals or the bell tower. If a line needs something not in the material, write it as '[NO SOURCE: what you need]'. No vocabulary like prophecy, fate, or gods.

[Verification request — Principle ④] Below the lines, write one line per voice_profile item marking met/violated, and flag any line you are not sure about with ★ plus one line of reasoning.

Note that all four principles are on this one page. The slots at the top are ①, the [Output format] block is ②, the [Hallucination guard] block is ③, and the [Verification request] block is ④. More than half the prompt is not "what to write" but constraints — how to write it, what not to write, how to verify itself. Given freedom, the model returns the average; the job is cutting the freedom down.

Step 3 — Claude's Output (Reconstructed from the Actual Session, Raw)

The vein is cooling. 12.4. Yesterday, 12.1. (detached)
Not a good sign. But my work is the record, nothing more. (cynical)
Why you ask is not my object of observation. (wary)

[Verification — voice_profile, 5 items]
- age_range (50s): Met. The restrained diction of an aging scholar.
- speech_pattern (numbers only): Met. Two figures in line 1. But no figures in lines 2 and 3 ★
- world_knowledge (nothing outside the guild): Met. No mention of outside affairs.
- taboos (no mysticism): Met. Zero uses of prophecy, fate, or gods.
- relationship (external variable): Met. Line 3 pushes the player outside the scope of observation.

★ Caution: read speech_pattern strictly as "speaks only in numbers" and
   lines 2 and 3 may be violations. Human call needed on whether to force a figure into every line.

The most valuable part of this output is not the three lines of dialogue but the single ★ line at the bottom. The AI flagged the ambiguous spot in its own output and handed it to a human. A good prompt lets the AI say "this part I am not sure about" — the direct effect of putting Principle ④ in.

Step 4 — Verification and Rejection (the Human's Seat)

The output is not accepted as is. A human rules on the ★ the AI raised. In this session, one line actually got caught.

In line 2, "Not a good sign," the adjective "good" collides with the voice_profile speech pattern ("almost never uses adjectives"). It is the exact spot the AI flagged with ★. Doren Vale is a character who speaks in figures instead of value-judgment adjectives, and "Not a good sign" slides into the diction of a stock old-man NPC. It is one line that blurs the tone.

So I re-request.

Line 2, "Not a good sign," uses an adjective ('good') and violates the voice_profile speech pattern.
Rewrite only that line in figures or observational vocabulary. Keep lines 1 and 3.
Format, hallucination, and verification rules apply unchanged.

The AI came back with line 2 as "Three years ago it was 9.0. That is my answer. (detached)" — the crisis shown through a change in figures, no adjectives, and the five voice_profile items passed again. One round trip closed it. Writing three tone-true lines by hand from scratch, versus one slot-filled prompt page + ★ review + one round trip — the latter carries the smaller review burden. That is this session's conclusion (based on my experience; absolute time varies with the difficulty of the NPC's tone, so read it as a direction).


22.1.3 The Four-Layer Structure — How to Stack a One-Page Prompt

Record on one page why the prompt above is stacked in that order, and every prompt after it becomes fill-in-the-blanks. Context stacks bottom to top, from heavy (rarely changes) to light (changes every time). The layers that do not change get cached to save cost (§22.1.5).

L0 Vision · Tone (world_premise · tone_manifesto) Rarely changes → cached. The foundation of Principle ① context. L1 voice_profile · naming rules · regional lore The 5 items that set this NPC apart + stated limits of knowledge → Principle ③ groundwork. L2 Adjacent lines · forbidden_names The scene's preceding lines · duplicate-banned names → keeps lines connected to their neighbors. L3 Task instruction (changes every time) [Output format]② · [Hallucination guard]③ · [Verification request]④ "Exactly 3 · 40 chars · (emotion) label · nothing beyond the material · 5-item self-check" Stack from bottom (heavy · cached) to top (light · swapped every call)

The one page in §22.1.2 is this diagram, exactly. L0 and L1 are pulled from existing material and pasted into slots (the foundation of Principles ① and ③), and L3 carries the three blocks — format, hallucination, verification (Principles ②, ③, ④). For the next NPC's dialogue, the only things that change are L1's voice_profile and L2's adjacent lines. The L0 and L3 skeleton is reused — which is how a prompt becomes a "library."


22.1.4 Prompts as Assets — The Library and Version Control

The npc_dialogue prompt above is not written once and thrown away. Prompts live in files, by domain and by task, and get called instead of rewritten each time. Project A's prompt folder looks like this.

prompts/
├── narrative/
│   ├── npc_dialogue_v3.txt        # ← the file from §22.1.2
│   ├── quest_synopsis_v2.txt
│   └── consistency_check_v1.txt
├── balance/
│   ├── change_proposal_v2.txt
│   └── outlier_analysis_v1.txt
├── content/
│   ├── city_npc_batch_v2.txt
│   └── side_quest_v3.txt
└── meta/
    ├── meeting_summary_v2.txt
    └── decision_card_v1.txt

The _v3 at the end of the filename is the point. A prompt is not finished once written — it is an asset close to a decision, so every change is measured for its effect on results before the version goes up. That is the path npc_dialogue took to v3.

flowchart LR
    A["Prompt v2
(no verification slot)"] --> B["Same N inputs through
both v2 and v3"] B --> C{"A/B compare
discard rate · tone violations · review time"} C -->|v3 lower| D["Adopt v3
npc_dialogue_v3.txt"] C -->|no difference| E["Keep v2
(change rejected)"] classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; classDef fail fill:#fee2e2,stroke:#dc2626,color:#7f1d1d; class B ai; class C human; class A data; class D pass; class E fail;

The actual change that took v2 to v3 is the [Verification request] block of §22.1.2. v2 had no slot for the AI to self-check its output item by item and attach ★. Adding that one block, the AI began reporting ambiguous lines first — as in §22.1.2 Step 4 — and the burden of a human reading everything from the top shrank. The version does not go up on "it feels better." The same input batch runs through v2 and v3 side by side, and only after confirming that tone-violation counts and review time actually fell does v3 get adopted.

The library's biggest payoff is the new team member. Call npc_dialogue_v3.txt on day one, and you start with the four-layer slot structure a senior refined over dozens of round trips. Before learning "how to write good prompts" in your bones, you already hold a well-written page in your hand.


22.1.5 Handling Cost Honestly — Caching and Caps

Longer prompts carry token costs. This chapter does not print unverified multipliers like "standardization erased a ×2 in cost." It speaks only of what can actually be measured.

Two structural devices hold cost down. First, caching is why §22.1.3 puts L0 and L1 at the bottom. Cache the vision-and-tone layer that rarely changes, and each call stops retransmitting — and re-billing — that layer. Pull NPC dialogue 100 times, and the gap between resending L0 100 times and caching it once widens with every call. Second, a per-call token cap keeps any one prompt from being stuffed with too much work at once.

What matters here is that the place where cost gets measured actually exists. Project A's atom system carries _economy_log/ (a token-and-time economy log) and _roi_report.md (ROI — return on investment — reporting) as operating metadata. The effect of prompt standardization is tracked in those logs as measurements, not asserted with plausible numbers in a table in this book. This book's principle is one of three:


22.1.6 Common Failures

Pattern Why It Fails Prescription
The one-liner "write 5 lines of dialogue" Zero context → generic-RPG average The four-layer slot prompt of §22.1.2
No "what they don't know" in voice_profile The AI invents lore beyond the material State the limits in the world_knowledge slot (Principle ③)
Output taken with no verification slot A human must read everything from the top Self-check + ★ via the [Verification request] block (Principle ④)
Prompts written fresh every time The same know-how rebuilt from zero The prompts/ library + versions
Prompt changes adopted on feel No way to confirm it got better A/B-measure the same inputs, then bump the version
Long context resent on every call Token cost stacks with the call count Cache L0·L1 + a per-call cap

The sixth is the last to be discovered. Cost does not hurt on a single call; it shows up in _economy_log after mass production has piled up.


Beyond Games. A one-line instruction calling up an average result — "fits anywhere, so it fits my job nowhere" — is not a problem unique to game dialogue. A prompt is the work order you hand a new hire, so put the four things on one page — what to look at when answering (context); count, length, banned items (output format); "invent nothing beyond the material" (hallucination guard); "mark for yourself which criteria you meet" (verification request) — and the output comes back cheaper to review. For example, when an HR manager requests a job-posting draft, stating "use only items in the role-requirements material; do not invent benefits or salary that are not in the material — mark them [NEEDS CONFIRMATION]" stops plausibly fabricated terms from leaking into the posting. Leave the one-page work order for a task you run often as a file, and it becomes a colleague's starting line.

22.1.7 Try It Yourself — One Step You Can Take Today

If you're solo, just this much: No library, no caching needed. Pick one NPC from your own game (or a game you love), write the five voice_profile items of §22.1.2 Step 1 by hand (age, speech pattern, scope of knowledge, taboos, relationship), paste in the Step 2 prompt body as is, and run it once. Out of the three lines you get, pick the one that clashes with the voice_profile and push back yourself — "this line violates the speech-pattern item; redo that line only" — and what each of the prompt's four slots does will sink in through your hands.

On a team, start with this one step. Pick one frequent task (say, NPC dialogue) and put a §22.1.2-style prompt page into prompts/narrative/ as a file. Check first that all four blocks made it in — slots, format, hallucination, verification — and that one file becomes the starting line for every new member of the team. Version control and caching come after.

Minimal web-chatbot path (no terminal) — This chapter's four principles work as is, with no files, library, or caching, in the single input box of a web chatbot (ChatGPT or Claude on the web). Prompt engineering is not a question of tools but of what goes on the one page. The two steps below are the main road. 1. Pick one task of your own and write down the five items of §22.1.2 Step 1 by hand (age, speech pattern, scope of knowledge, taboos, relationship — outside games, read them as "audience, register, scope of evidence, bans, relationship"). No YAML, no files; writing them in the chatbot's input box is enough. 2. Below that, paste in the prompt body of §22.1.2 Step 2 as is, checking only that all four blocks are present — [Output format] (count, length, labels, bans), [Hallucination guard] ("invent nothing beyond the material; if needed, [NO SOURCE]"), [Verification request] ("mark met/violated per item; if unsure, ★"). Run it, have a human rule on only the ★-flagged lines, push back once, and one cycle closes. Library, versions, and caching come in only once you find yourself reusing the same prompt.


22.1.8 A Preview of the Next Chapter

22.2 covers hallucination and safety. If this chapter's Principle ③ (the hallucination-guard slot) is the first line of defense inside the single prompt page, 22.2 looks at the layered defenses that catch, at the operations level, the hallucinations that break through.


Key Takeaways

Next Chapter Preview

22.2 The Colleague Who Lies with Confidence — Stopping Hallucinations with a Verification Gate

Primary audience: game designers producing documents, data, and decision records at volume with AI (mid-sized teams of 10–50) Scaled-down version for solo/hobbyist readers: §22.2.7 "If You're Solo, Just This Much"

It happened on a day I had AI summarize 17 sets of meeting notes into decision cards. The output was clean. Decision IDs, cited meeting dates, even a one-line rationale — the format was flawless. One of the cards read, "Cooldown policy finalized at the combat task force meeting on 2026-04-18." The problem: there was no combat task force meeting that day. The AI had blended the agenda and date of different meetings into one plausible card, and because the format was flawless, it very nearly went straight into the team's decision record.

This is hallucination. The less an LLM knows, the more confidently it answers. In human terms, it is the colleague who states flatly in a meeting, "Oh, that was already decided," when no such decision was ever made. When that one remark flows into a data sheet, a customer support reply, or an atom asset, it becomes an incident. This chapter is not about how to silence that colleague — that is impossible — but about how to build a verification gate that everything he says must pass through before it is accepted (think of it as a quality gate for AI output). General theory of hallucination fills other books, so this chapter focuses only on the place where an AI workflow stops it.


22.2.1 Hallucinations Never Hit Zero — So We Put Up a 'Gate'

No prompt makes hallucinations zero. Bigger models and better prompts lower the frequency, but never to 0. So the starting point of operations must be not "eliminate hallucinations" but "put a gate in place that catches them before they reach decisions and data."

The gate has one core principle. Anything an LLM can fabricate (citations, numbers, IDs) gets verified somewhere that is not an LLM. The source of verification is one of three: code (deterministic), the original document (grep), or human eyes. Asking the LLM again — "check whether this is right" — can serve as one stage of the gate, but it is an assist, not the final judge.

Let me first sort out the point game designers most often get confused about. The areas where hallucination comes easily and the areas where it does not are clearly different.

Task Hallucination Risk Why Gate
Numeric calculation (rewards, probabilities) Very high LLMs estimate arithmetic Calculations go to code; ban them for the LLM
Citations (meetings, decision IDs) High Fabricates plausible nonexistent sources grep check against the original
Classification (tags, categories) Medium Mixes up labels Deterministic comparison possible
Summarization and inference Medium Adds or drops items Self-verification + human gate
Creative writing (flavor text) Low No ground truth, so 'hallucination' barely applies Tone review gate

The first row is the simplest prescription. Don't give numbers to the LLM. Hand them to a deterministic tool, the way you hand multiplication to a calculator. The second row (citations) is the spine of this chapter. In tasks where an original exists and the LLM is restating it — meeting-note summaries, decision cards — hallucination is at its most dangerous and also at its most catchable, because the original gives you something to check against.


22.2.2 [Worked Transcript] Catching Hallucinations in a Meeting-Note Summary with Self-Verification

Writing "verify it" in the abstract says nothing about what to do or how. We will follow one full cycle — summarizing one set of meeting notes and catching the hallucination in that summary — from input all the way to the re-request. The prompts below can be copied as they are; the outputs are reconstructions of an actual session.

Step 1 — Input: Feed in the Original Meeting Notes as They Are

First, there is an original to summarize. This becomes the reference point for verification. Let the LLM summarize from "memory" without an original, and there is nothing left to check against.

# Meeting Notes — 2026-04-15 Combat Balance Review

Attendees: Minsoo Lee, Team Member A, Team Member B
Agenda:
1. AoE skill cooldown — opinion that the current 8s is too short. Reviewing a raise to 12s.
   Conclusion: tentatively decided at 12s; finalize after telemetry check in the next build.
2. Boss phase-transition HP — proposal to lower from 50% to 30%.
   Conclusion: on hold. Conflicts with the art presentation schedule; rediscuss next week.
3. Duration of new status effect 'Freeze' — discussed, but no value set.

Step 2 — The First-Pass Summary Prompt

Summarize the attached meeting notes into decision cards.

Rules:
1) One card per agenda item. Format: [Agenda | Conclusion | Status (Final/Tentative/On hold) | Rationale]
2) Do not generate information that is not in the notes. Do not guess numbers, dates, or attendees.
3) For items with no conclusion, set the status to 'Undecided' and do not invent one.

Output format: table only. No commentary.

Step 3 — Claude's First-Pass Output (Reconstructed, Excerpted)

| Agenda                | Conclusion        | Status    | Rationale                                |
|-----------------------|-------------------|-----------|------------------------------------------|
| AoE skill cooldown    | 8s → 12s          | Tentative | Finalize after telemetry check           |
| Boss phase transition | 50% → 30% on hold | On hold   | Conflict with art presentation schedule  |
| Freeze duration       | Set to 3s         | Final     | Standard status-effect duration applied  |

Look at the third card. The original says only that "Freeze duration was discussed, but no value set." The output, however, wrote "set to 3s / Final / standard baseline applied." All three fields are information that is not in the meeting notes. The LLM filled the blanks from general knowledge — "a status effect usually lasts a few seconds." The format is flawless, so a casual look catches nothing.

Step 4 — The Self-Verification Re-Request (the Heart of the Gate)

Here, instead of accepting the output as is, force a self-verification that checks it against the original, item by item. This is the key move of the chapter.

Check the 3 decision cards you just made against the attached original meeting notes, one item at a time.

For every field of every card (conclusion, status, rationale, numbers):
- Is it written directly in the original? → [match]
- Is it something you added that is not in the original? → mark it [generated] and state that it has no source
- Is anything in the original missing from the card? → report it as [omission]

Revert every [generated] field with no basis in the original to 'Undecided'.

Step 5 — Claude's Self-Verification Output (Reconstructed, Raw)

Comparison results:

Card 1 (AoE skill cooldown)
- Conclusion "8s → 12s": [match] agenda item 1, "raise to 12s"
- Status "Tentative": [match] "tentatively decided"
- Rationale "finalize after telemetry check": [match]

Card 2 (boss phase transition)
- Conclusion "50% → 30% on hold": [match]
- Status "On hold": [match] "on hold. Rediscuss next week"
- Rationale "conflict with art presentation schedule": [match]
- [omission] The original's "rediscuss next week" condition is missing from the card. Needs to be added to the rationale.

Card 3 (Freeze duration)
- Conclusion "set to 3s": [generated] — the original says "no value set". The 3s is a value I added.
- Status "Final": [generated] — the original reaches no conclusion.
- Rationale "standard status-effect duration applied": [generated] — no such baseline is mentioned in the original.
  → Per the rules, I am reverting card 3 as follows:
  | Freeze duration | (value not set) | Undecided | Discussed in the meeting, no conclusion |

This output is the most valuable part of the chapter. The AI checked its own output against the original, confessed "the 3s is a value I added," and reverted card 3 to 'Undecided' per the rules. At the same time it caught an omission in card 2 (the "rediscuss next week" condition) that the human had missed too. Hallucination (adding what isn't there) and omission (dropping what is) are two sides of the same coin, and the same comparison catches both.

The caveat is just as clear. This self-verification is not a cure-all. If the LLM misreads the original, it can deliver a wrong comparison result with full confidence. So self-verification is the first stage of the gate, and when the original is short, a human backs it up once more with grep. Self-verification catches blatant generation like card 3 almost every time, but subtle paraphrase and nuance distortion ultimately fall to the human gate at the end.


22.2.3 The Verification Gate — One Flowchart

Generalizing the cycle above, the gates an AI output passes through before it reaches decisions and data look like this. Human hands touch only two places: the very front, where a clean original goes in, and the very end, where a person makes the judgments the automated gates cannot.

flowchart TB
    A["Original (meeting notes, data sheets)
The reference point for verification"] --> B["AI first-pass generation
Summaries, decision cards, classification"] B --> C{"Self-verification
Check against original item by item
[match]/[generated]/[omission]"} C -->|Generated field found| D["Generated fields → reverted to 'Undecided'"] D --> E C -->|Match| E{"Deterministic gate
grep check on numbers, IDs, citations"} E -->|Number/ID mismatch| F["Reject + re-request
(correct to original values)"] F --> B E -->|Pass| G["Human gate
Judgment on nuance, tone, context"] G -->|Rejected| F G -->|Approved| H["Decision record / applied to build"] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; class E code; class B,C ai; class G human; class A data; class H pass;

The gate is threefold because each stage catches something different. Self-verification has the LLM check itself for whether it added what isn't there; the deterministic gate uses code to check whether numbers and IDs match the original character for character; the human gate catches what is technically right but contextually off. Turn on only one stage, and incidents leak through where the other two used to stand. In §22.2.2, the "3s" on card 3 is caught at the first stage (self-verification), card 2's omission at the first stage, and if self-verification had misread "12s" as "21s", the second stage (grep) would catch it.


22.2.4 The Gate Can Fail, but It Never Blocks the Flow — Hook Safety Design

When you put a verification gate into an automated pipeline, there is one incident beginners cause most often: building it so that when the gate itself crashes, all work stops. If grep dies on an encoding error or the manifest file is corrupted, the code that was supposed to help with verification ends up blocking the user's work outright. Within a week or two the team says, "let's turn that verification off."

Let me quote, as is, how the JIT atom-injection hook actually in operation in this book (inject_memory.py) handles this problem. This hook cuts in every time the user types a prompt and injects relevant memory — an always-on gate, so to speak. One line is spelled out in its design-principles comment.

Design principles:
- Always exit 0 (even on failure, never disrupt the user's flow)
- No match → empty response (normal)

And the principle is implemented consistently throughout the code. If stdin parsing fails, if the manifest JSON is corrupted, if reading an atom body fails — everything falls through to emit_empty() and exit 0.

def emit_empty() -> None:
    sys.exit(0)

def main() -> None:
    try:
        ...
        payload = json.loads(raw)
    except Exception:
        emit_empty()        # even broken input passes through quietly
        return
    ...
    try:
        manifest = json.loads(MANIFEST_PATH.read_text(encoding="utf-8"))
    except Exception:
        emit_empty()        # a broken manifest never blocks the work
        return

if __name__ == "__main__":
    try:
        main()
    except Exception:
        emit_empty()        # the last net for any exception

The heart of the design is that it separates gate failure from content failure. When the hook fails to inject memory, the user just gets "an ordinary session with no memory attached" — not an incident that blocks work. A verification gate must work the same way. If the grep gate can't run because of an encoding problem, you don't pass the card through — you mark it "automated verification failed — to the human gate" and hand it to the human stage. A dead gate must not auto-approve unverified output, and at the same time a dead gate must not halt the whole pipeline. The safe default that satisfies both is "quietly hand it to a human." The except: emit_empty() in inject_memory.py is the minimal implementation of exactly that pattern.


22.2.5 Handling Hallucination Rates Honestly

The temptation to put a table like "we cut the hallucination rate from 89% to 3%" into this chapter is strong. Numbers like that, without a disclosed measurement method, erode the book's credibility. This book holds to one of three rules.

First, state numbers only for what can be measured. To promise a hallucination rate, you have to define the denominator and the numerator. The denominator is "the number of decision cards reviewed"; the numerator is "the number of cards where the comparison against the original caught at least one [generated]/[omission]." Without this definition, "a 5% hallucination rate" is hollow. This is exactly how I counted while reviewing meeting-note summaries in the early days of adoption, and that sample is small — a directional value, not a precise population parameter.

Second, state model-to-model comparisons as direction only. The direction "bigger models hallucinate less than smaller ones" is observed consistently. But absolute figures like "Opus 3%, an open 7B model 20%" swing widely with task, prompt, and domain, so this book claims no absolute values. We take only the direction (the bigger the model, the fewer — traded off against cost).

Third, cite public standards as they are. This chapter has almost no standard figures to make up, but settings like temperature are public facts in model API documentation. Verification and analysis tasks run with temperature low (closer to deterministic), creative tasks with it high — that is not an estimate but the definition of how the API behaves.

So the measurable indicators this chapter actually promises are three — the count of [generated] detections (hallucinations caught by self-verification), the count of grep-gate rejections (number/ID mismatches), and the count of human-gate rejections. All three can be counted from logs every quarter, and in meetings you can speak in numbers instead of "feel."


22.2.6 Common Failures

Pattern Why It Fails Prescription
Accepting AI summaries on format alone Hallucination is hardest to catch when the format is flawless Self-verification against the original (§22.2.2, step 4)
Summarizing from LLM memory without the original No reference point to check against, so unverifiable Put the original into the input first
Leaving numeric calculation to the LLM Arithmetic gets estimated, differently every time Calculations go to deterministic tools (§22.2.1)
All work stops when the verification gate crashes The team turns the gate off exit 0 + hand off to the human stage (§22.2.4)
Auto-approving unverified output when the gate dies Hallucinations pass straight through Gate failure = mark as 'unverified'
Trusting self-verification as the final verdict An LLM that misreads the original misjudges with confidence For short originals, run human grep in parallel

Beyond Games. The "colleague who lies with confidence" — an AI that fabricates nonexistent meeting dates and decisions in flawless format — is just as dangerous in every document summary, not only in game decision cards. Hallucination is hardest to catch when the format is flawless, so for any task that has an original (meeting-note summaries, contract excerpts, report digests), the key is not to accept the output as is but to force self-verification: "check against the original item by item, and mark anything you added as [generated]." For example, have a legal assistant summarize a contract and then make the AI compare amounts, dates, and clause numbers against the original character for character — the AI confesses "this penalty figure is a value I added" and reverts the blank to 'undecided.' Don't hand numeric calculation to AI at all — give it to a calculator or a formula — and design automated verification tools so that even when they fail, they don't block the work but hand off as "unverified — human check."

22.2.7 Try It Yourself — One Step You Can Take Today

If You're Solo, Just This Much: No code, no hooks needed. Have AI summarize one short document of your own (a meeting memo, patch notes, one page of a design doc), then paste in the step-4 self-verification prompt from §22.2.2 as it is. One line — "check against the original item by item, and mark anything you added as [generated]" — and the AI starts reporting its own hallucinations. Get even one [generated] confession, and you will feel, hands-on, why AI summaries must never be taken on faith.

If you're on a team, start with this one step. Pin the self-verification stage into the default prompt for the decision cards and summaries AI produces (§22.2.2). Next, pick only the fields that must match the original character for character — numbers, decision IDs, dates — and turn the grep comparison into code. That verification code must be designed, like inject_memory.py, so that failure never blocks the work (exit 0 + an 'unverified' mark) (§22.2.4). With just those two stages — self-verification and grep — you head off the most common incident first: flawlessly formatted hallucinations seeping into the decision record.


Key Takeaways

Next Chapter Preview

22.3 AI Cost Management — Enforcing the Token Budget in Code

Primary audience: design leads who roll out AI tools to their team and own the cost (mid-size teams of 10–50) Scaled-down version for solo/hobbyist readers: §22.3.9, "If You're Solo, Just This Much"

If a cost chapter cites fake costs, it contradicts itself. So this chapter does not build a polished table of "how much my team saved per month." Instead it uses only two kinds of numbers. One is the published token pricing anyone can verify (per-1M-token rates by model); the other is the set of constants pinned in the hook code I actually run (max_atom_body = 6000, max_matches = 3). Neither is made up — both are quoted.

What makes AI cost scary is not the size of the bill. It's that you can't see it. In the first month of adoption, calls are few and the bill is small. Then contexts get longer and calls more frequent, and one quarter the bill gains a digit. Here is this chapter's conclusion up front — you control cost not with a resolution to "use it sparingly" but with code that forcibly trims tokens on every call. The wrapper and the truncate hold the line, not human willpower.


22.3.1 LLM Cost Is Effectively One Line Item: Input Tokens

There are four billing items — input, output, cache reads, and cache writes — but in practice the bill is dominated by input tokens. The reason is simple: almost every task where a game designer uses AI takes the form of "feed in a long context, get back a short answer." Cram in the L0 vision document, the atom library, adjacent city pages, and data sheet excerpts, and the input runs to tens of thousands of tokens — while the output is a single table, a few hundred tokens.

So the first priority of cost control is not "shrink the output" but "where do we cut input tokens." That one line drives the rest of this chapter.

Let me pin down the published per-model rates first. The table below shows the rates Anthropic publishes per 1M (one million) tokens — a snapshot quoting the published pricing of the generation current as of this writing (the then-latest tier of Opus, Sonnet, and Haiku) (quoted from the official published price list — rates change with model generation and over time, so check the current price sheet before applying any of this). As Appendix K lays out, what does not change here is not the absolute rates but the ratio between the three tiers. So read the table below not as "today's bill" but as a structure: each tier you step down drops the unit rate by close to an order of magnitude.

Model Input per 1M tokens Output per 1M tokens Notes
Claude Opus $15 $75 Top-tier reasoning (published rate)
Claude Sonnet $3 $15 Mid-tier — input rate 1/5 of Opus
Claude Haiku $0.80 $4 Lightweight — input rate about 1/19 of Opus
Cache hit (read) about 1/10 of the standard input rate When cached input is reused (published caching policy)

The last two rows are the point. Run the same task on Haiku instead of Opus and the input-token rate is about 1/19; route the same context through the cache and that portion of the input bills at about 1/10. The two big levers of cost reduction come from here — model right-sizing and caching. Neither is "use less"; both are "do the same work at a cheaper rate."

Savings come from rate differences, not willpower. Dropping Opus to Haiku cuts about 19x; riding the cache cuts about 10x — automatically.


22.3.2 The Biggest Input Cost Is the Context Injected on Every Call

There is a cost that accumulates more quietly than per-task rates: the context automatically attached to every call. On my personal PC, a hook runs that inserts relevant memory (atoms) every time I type a prompt (the UserPromptSubmit hook, inject_memory.py). It's a convenience feature — and at the same time the number-one suspect for cost leakage. Long atom bodies enter the context with every input, so left uncontrolled, input tokens balloon call after call.

So this hook has three layers of cost-cutting safeguards pinned in place. Not abstractions — constants in the actual code.

# inject_memory.py — UserPromptSubmit hook (actual production code, excerpt)
# Design principles (from the docstring):
#   - Always exit 0 (never block the user's flow, even on failure)
#   - Inject at most 3 atoms, in descending score order
#   - Truncate any atom body over 6000 characters

# (1) Read the budget constants from the manifest config
max_matches = cfg.get("max_matches", 3)      # max atoms per call
max_body    = cfg.get("max_atom_body", 6000) # body cap per atom (chars)

# (2) Sort by score, descending — fill the expensive slots by value
atoms_sorted = sorted(atoms, key=lambda a: a.get("score", 0), reverse=True)

matches = []
for atom in atoms_sorted:
    if len(matches) >= max_matches:   # (Guard A) cut off at 3
        break
    if re.search(atom["regex"], prompt, re.IGNORECASE):
        matches.append(atom)

# (3) Cut the body at 6000 characters on injection
for atom in matches:
    body = atom_path.read_text(encoding="utf-8")
    if len(body) > max_body:          # (Guard B) truncate
        body = body[:max_body] + "\n\n[...truncated]\n"

All three cost guards are right there.

These three constants are exactly the per-call ceiling on input tokens. As a rough estimate, one 6,000-character atom is on the order of several thousand tokens in Korean (exact token counts vary by tokenizer and language, so read this as "a ceiling exists," not as an absolute value). 3 × 6,000 characters is the injection budget per call, and the code cuts everything above it. No human has to spot "too many atoms got attached" by eye.


22.3.3 [Worked Transcript] How One 6000-Character Truncate Line Holds Down Cost

Saying "the truncate holds down cost" in words is hollow. When I set this constant, I ran a full cycle with the AI from start to finish. Below is a faithful reproduction of that session. The input prompts can be copied verbatim; the outputs are reconstructed from the actual session.

Step 1 — Input: Throw the Problem at It As Is

Right after the hook first went live, _injection_log.txt recorded a call where an atom body had been injected whole. If long atoms (retrospectives, health context, and so on) go in at full size on every call, input tokens accumulate fastest on the tasks you use most. I asked the AI where to cut.

Step 2 — Prompt: Force Limits and Format

I have a hook (inject_memory.py) that auto-injects atoms into the context on
every call. I want a truncate cap to hold down input token cost — propose a
cap on body characters per atom and a cap on atoms injected per call. Too
short and the core gets cut; too long and tokens leak, so spell out that
trade-off too, and flag anything you're not sure about for a human to decide.
(Intent: the caps will be externalized to config.) As a table:
item | proposed value | rationale | confidence.

Note that half the prompt is "forcing limits and format." Give an AI free rein on a cost question and you get an evasive "it depends." The core of this prompt is to force concrete numbers, make the trade-offs explicit, and leave an escape hatch: report what you're not sure of.

Step 3 — Claude's Output (Reconstructed, Excerpt)

| Item            | Proposed   | Rationale                                       | Confidence |
|-----------------|------------|-------------------------------------------------|------------|
| max_atom_body   | 6000 chars | The 'core + evidence' of one Korean atom mostly | Medium     |
|                 |            | fits in 4–6k chars. At 6000 the head (gist,     |            |
|                 |            | rules) is nearly intact; only the tail          |            |
|                 |            | (example lists) gets cut.                       |            |
| max_matches     | 3          | Topics that strongly match one input are        | High       |
|                 |            | usually 1–3. 4+ signals the regex is too broad. |            |
| Externalize to  | _jit_manifest.json > config | Tune live without redeploying code | High |
| Truncate marker | append [...truncated] tail  | Model and humans see it was cut    | Medium |
| Not confident   | whether 6000 is exact       | Char-to-token conversion varies by | ★Low |
|                 | (char→token mapping uncertain) | language and model — chars are approximate | |

The most valuable part of the output is not the table — it's the ★ row at the bottom. The AI reported, on its own, that it doesn't know whether 6000 is the exact optimum, because char-to-token conversion is uncertain. A good prompt makes the AI say it doesn't know.

Step 4 — Verify and Reject (the Human's Seat)

I did not accept this output as is. The AI presented 6000 characters as "the line where the core survives," but you have to see for yourself what actually gets cut at 6000. I picked the longest atoms (health context, accumulated retrospectives) and eyeballed the 6000-character mark.

One case got caught. The 6000-character mark severed a table row in the middle of a body, leaving a line like | Vitamin D | 1000 IU | chopped in half. The meaning survived, but a broken table row going to the model as malformed Markdown bothered me. So I asked again.

Step 5 — Follow-Up Request

When cutting at 6000 characters, a Markdown table row sometimes gets severed
mid-row. Would snapping the cut to the previous newline boundary be better? The
fix must not outgrow the truncate itself — weigh cost vs. value and decide.

The AI answered: "Line-boundary snapping is one line — body[:max_body].rsplit('\n', 1)[0] — but the purpose of the truncate is a token ceiling, not table integrity, so one broken table row has negligible impact on cost or behavior. The [...truncated] marker already lets the model recognize 'it was cut here,' so I recommend not doing the line snap. If you must, the one-line fix is enough."

That answer was right. If cost-control code also chases table integrity, the guard code itself bloats. The production code keeps the simple cut: body[:max_body] + "[...truncated]". The AI's first proposal (6000 characters) verified by a human, and the human's urge to over-engineer pushed back by the AI — a cycle closed in one round trip.


22.3.4 The Cost Guard Structure — At a Glance

Here is the full flow, recorded as a diagram: how the constants set in that session cut input tokens on a real call.

flowchart TB
    A["User enters a prompt"] --> B["regex matching
iterate the 17-atom library"] B --> C{"Guard A
len(matches) >= 3?"} C -->|"blocked from the 4th on"| D["break — not injected"] C -->|"slots available"| E["Fill slots in
descending score order"] E --> F{"Guard B
body > 6000 chars?"} F -->|"over"| G["body[:6000] + [...truncated]"] F -->|"within"| H["Inject in full"] G --> I["Context injection
(input tokens: at or below ceiling, guaranteed)"] H --> I D --> I classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; classDef fail fill:#fee2e2,stroke:#dc2626,color:#7f1d1d; class B,C,E,F,G,H code; class A human; class I pass; class D fail;

The point of this diagram: whatever the user types, there is a ceiling on injected tokens per call. The ceiling is 3 × 6000 chars (plus the marker), and the code unconditionally cuts everything above it. Cost does not lean on the user's self-restraint. Guards A and B fire mechanically on every call.

The same philosophy repeats at the tool level. My company system has a policy that pins the number of wrapper skills exposed in the global slots to exactly 12 (atom skill_listing_budget_wrapper_only_policy). If the global * wrapper count is not 12 at session start, a cleanup script runs automatically. Nominally it's "slot tidiness," but in essence it's session-start token budget protection — the cost of the skill list landing in context is bounded to 12 entries' worth. The 3-atom injection cap and the 12-skill exposure cap are the same idea in two applications.


22.3.5 Model Allocation by Task — 80% at a Cheaper Rate

Guards cap tokens per call; model choice sets the rate those tokens bill at. In the §22.3.1 table, input rates ran Opus:Sonnet:Haiku ≈ 19:4:1. Running every task on Opus therefore means paying a 19x rate even on simple jobs like classification and substitution.

Allocate rates by task complexity.

Task type Recommended model Why
Verification, legal-adjacent work, decision analysis Opus Tasks where a mistake is a big incident — don't skimp on the rate
Reports, summaries, natural-language polishing Sonnet Quality matters, but top-tier reasoning is unnecessary
Classification, tagging, keyword extraction Haiku Simple patterns — about 1/19 of the Opus rate is plenty
Simple mapping and substitution Haiku, or deterministic code Often no LLM is needed at all

In my experience, most tasks are fine on Sonnet or Haiku. Expensive models are for tasks that are expensive to get wrong. One trap, though — go too cheap and hallucinations rise, and the verification cost eats the savings (this ties directly into §22.2, the previous chapter on hallucination and safety). So model allocation is not "cheapest always"; it's a split — cheap where being wrong is cheap, expensive where being wrong is expensive.

The last row, "simple mapping/substitution → deterministic," is often the biggest saving. Work with a single fixed answer — name substitution, rule-based mapping — has no need to call an LLM. Driving the call itself to zero is the cheapest call there is.


22.3.6 Caching — The Same Input at 1/10 the Rate

Even with tokens capped per call (guards) and rates lowered (model allocation), cost still leaks if the same context is re-sent fresh on every call. Long inputs that barely change — the L0 vision document, the atom library, the per-discipline style guide — get cached. On a cache hit, that portion of the input bills at about 1/10 of the standard rate (§22.3.1 table).

# Mark unchanging context with cache_control — about 1/10 the rate on cache hits
messages = [
    {"role": "system", "content": SYSTEM_PROMPT},
    {"role": "user", "content": [
        {"type": "text", "text": L0_VISION,    "cache_control": {"type": "ephemeral"}},
        {"type": "text", "text": ATOM_LIBRARY, "cache_control": {"type": "ephemeral"}},
        {"type": "text", "text": SPECIFIC_TASK},  # only the part that changes each time stays outside the cache
    ]},
]

The key is to separate what changes from what doesn't. The cache only hits when the front of the input is identical, so the fixed context (L0, atoms) goes first and the per-call task instruction goes last.

What to put on the cache is decided by change frequency.

Context Cache? Why
L0 vision (nearly immutable) Good fit Changes only every few days to weeks
atom library Good fit Updated only at retrospectives
Discipline style guide Good fit Changes on a quarterly cadence
Recent meeting notes Poor fit Changes daily — low hit rate
User input Poor fit Unique to every call

Cache TTLs can be as short as a few minutes, so the payoff is biggest for work that hits the same context back to back (like mass-producing 30 cities, reusing the same L0 thirty times). On one-off questions you pay the cache-write cost and never get a hit — it can be a net loss. So apply caching selectively, to work that uses the same context often and consecutively.


22.3.7 How to Handle Numbers Honestly

A cost chapter is where the temptation is strongest to drop in a table like "we cut $5,000 a month down to $1,000." Absolute savings vary wildly with team size and workload, and the moment you invent one, the cost chapter is lying about cost — a self-contradiction. This chapter used only three kinds of numbers.

First, published rates are quoted as is. The §22.3.1 figures — Opus $15 / Sonnet $3 / Haiku $0.80 (input, per 1M tokens) and the cache-hit rate of about 1/10 — are Anthropic's published pricing. The 19:4:1 input-rate ratio and the roughly 10x caching saving fall out of that published pricing by arithmetic — calculation, not estimation.

Second, code constants quote the code. max_atom_body = 6000 and max_matches = 3 are values recorded in the actual inject_memory.py and _jit_manifest.json. Not metaphors — real files.

Third, what I don't know, I say I don't know. "How many tokens is 6000 characters" varies by tokenizer, language, and model, so character counts are approximations. In §22.3.3, the AI flagged the same point with a ★. That is why nowhere in this chapter will you find a conversion table like "6000 chars = N tokens = $X saved." Instead of absolute savings, it speaks only in direction and ratios (19x, 10x).

Every cost figure in this chapter is either published pricing (Anthropic's rate sheet), a constant pinned in code (inject_memory.py, _jit_manifest.json), or an approximation explicitly marked "unknown."


22.3.8 Common Failures

Pattern Why it fails Fix
Top-tier model for every task About 19x the rate even on classification and substitution Model allocation by task (§22.3.5)
Re-sending the same context every call Throws away the 1/10 cache-hit rate Cache the fixed context (§22.3.6)
No cap on auto-injection The whole atom library injected on every call Count and length guards (§22.3.2)
Managing cost with a "use it sparingly" resolution Human restraint can't stop a surge Pin the ceiling as a code constant
Calling an LLM for deterministic work The cheapest call is "no call" Split mapping/substitution out into code

The fourth row is the heart of it. Leave cost control to human willpower and it will leak — without fail. Willpower is the first thing to collapse when you're busy, and cost grows fastest when you're busy. That's why the control has to be a code constant like max_matches = 3.


Beyond Games. What makes AI cost scary is not the size of the bill but its invisibility — and that's the same whether you're a game team or a marketing team. You contain cost with structure, not with a resolution to spend less. First, allocate model rates to task difficulty — running simple classification and tagging on the top model means paying several times the rate for the same work, and for simple mapping and substitution, not calling at all (handle it with rules and formulas) is the cheapest call. Second, cache long inputs that rarely change (company overviews, policy documents, glossaries) to cut re-transmission cost. For example, classifying customer inquiries is fine on a lightweight model, and only complex contract review goes to the top model — you protect quality while splitting the rates. If there's a spot where long context attaches automatically, set a cap on how many items and how much length can attach at once — that structurally prevents the day the bill gains a digit.

22.3.9 Try It Yourself — One Step You Can Take Today

If You're Solo, Just This Much: you don't need a hook or a manifest. In the AI tool you use most, drop the model one tier for your next single task (the summary you ran on Opus → Sonnet; the Sonnet classification → Haiku). If the output quality holds, that task is permanently pinned at the cheaper rate. Just asking "does this task really need the top model?" once per task gets you half the savings.

If you're on a team, start with this one step. Find one place that injects context automatically (a hook, a system prompt, RAG) and put the two guards from §22.3.2 into its code — one cap on injection count, one cap on body length. Externalize the caps to config, the way inject_memory.py does, and you can tune the numbers in production without redeploying code. Two lines of guard structurally prevent the accident where "one day the bill gains a digit."

Summed up as setup → prompt → verify — setup: put count and length cap constants at the auto-injection point and pull them out into config. prompt: have the AI propose the cap values in the §22.3.3 format, forcing trade-offs and confidence levels. verify: pick your longest input and check with your own eyes what gets cut at the cap.


Key Takeaways

Next Chapter Preview

22.4 Copyright and Ethics — Closing an Output's Rights, Disclosure, and Agreement in One Procedure

Primary readers: game directors and leads responsible for AI adoption (mid-size teams of 10–50) Scaled-down version for solo/hobbyist readers: §22.4.9 "If You're Solo, Just This Much"

Two months before launch, a meeting once ground to a halt over a single city illustration a concept artist had made. Someone asked, "This was generated with AI, right? So is the copyright ours — or can it not even be registered?" Nobody could answer. Three opinions came out of that room: "The AI made it, so it isn't ours." "We paid to run the model, so it's ours." "There's no law yet, so just use it." All three were wrong. And the question was never just a legal-affairs issue. What the role of the artist who made that illustration actually was, and how the team had agreed on AI use, were all hanging on it at once.

This chapter does not treat copyright and ethics separately, because in practice they are the front and back of the same question. "Who owns the rights to this output" (copyright) reduces directly to "how much did a human intervene in this output" (ethics and roles). The registration requirement the Korea Copyright Commission nailed down in 2025 sits exactly at that point. So the spine of this chapter is a single worked transcript — it actually judges the copyright registrability of one piece of AI concept art and follows, from input to decision, how that judgment leads into the team's role agreement.

Author's Actual Operations Note The design_intent_vs_automation_boundary atom and the _economy_log / _roi_report.md cited in this chapter are anonymized versions of governance assets I actually operate at work. The atom names and log file names are the real operating names (only company- and project-specific identifiers were replaced for IP protection). The worked transcript's output is a reconstruction of an actual judgment session.


22.4.1 Authority Comes from Public Guidelines, Not from "Feel"

Many books simply write that AI copyright is "still murky because the law isn't there yet." That's only half right. In June 2025, when the Ministry of Culture, Sports and Tourism and the Korea Copyright Commission published the "Guide to Copyright Registration for Works Made Using Generative AI," the line for registrability became clear — at least in Korea. There is no need to make anything up.

The guide's core compresses into one sentence: the requirement for copyright registration is "human creative contribution." From there, two categories split.

Category Definition Registration
GAI output A result the AI produced with no human creative contribution Not possible
GAI-assisted work The parts of a result made using AI as a tool where creative contribution is recognized Possible

The guide then lays out three paths to recognition as an "assisted work": ① the user fed their own copyrighted work in as a prompt and its creativity shows in the output; ② the additional work of modifying or augmenting the output is creative; ③ the selection, arrangement, or composition of outputs is creative. The two axes of judgment are "controllability" and "predictability." Creativity is recognized when the creator clearly decides what they want to express and can pull the result toward that intent.

This part is decisive. What the guide says in legal language — "controllability and predictability" — is the same thing this book has repeated since §1.1: "the designer provides intent" (the planner_provides_intent_not_recommendation atom). An output handed wholesale to the AI has no control and no prediction, so it has no copyright either; an output where a person input the intent, then reviewed and reconstructed the result, carries rights with it. Copyright registrability and the conditions of a good AI workflow sit on the same line.

There is one more public standard. The AI Framework Act — Korea's basic law on AI, taking effect in 2026 — imposes a transparency duty (disclosure of the fact of AI generation) on generative AI outputs. Registration (the side that claims rights) and disclosure (the side that states the fact of use) are separate duties. Whether rights arise or not, the fact that you used AI must be disclosed. These two public standards become the primary input for the rulebook we hand the AI in this chapter.


22.4.2 [Worked Transcript] Judging the Registrability of One Piece of AI Concept Art

Back to the illustration from the opening. Instead of judging it by "feel," I feed the §22.4.1 guide's criteria to the AI as a rulebook and have it do the first-pass classification. The human makes only the final call. The input prompt below can be copied and used as is; the output is a reconstruction of an actual judgment session.

Step 1 — Input: Hand Over the Output's Generation History as Is

The input for the judgment is not the illustration but the log of how that illustration was made. This already lives in the asset metadata, so all you do is extract it.

# asset_concept_city021_v4.meta.yaml — generation history of the output under judgment
asset_id: concept_city021_v4
asset_type: concept_illustration
created_by: Team member A (concept artist)
generation_log:
  - step: 1
    actor: Team member A
    action: "Attached a hand-drawn rough sketch of the city layout as the input image"
  - step: 2
    actor: AI (image_model)
    action: "Generated 4 variations from the rough plus a prompt"
    prompt: "stone observatory tower over sealed magic core, cold arid, scholar guild, muted palette"
  - step: 3
    actor: Team member A
    action: "Picked 1 of the 4; retouched the bell tower silhouette, saturation, and composition by hand (about 40% of the area reworked)"
  - step: 4
    actor: Team member A
    action: "Designed the background sealing-array sigil by hand and composited it in"
ai_generated_disclosure: true   # satisfies the AI Framework Act disclosure duty

Step 2 — Prompt: Classify by the Guide's Criteria, but Force the Reasoning

The attached meta.yaml is the generation history of one concept illustration. Using the Korea
Copyright Commission's "Guide to Copyright Registration for Works Made Using Generative AI" (2025),
classify its registrability as one of: registrable / partially registrable / not registrable.
Give one line of reasoning for each of the guide's two axes (controllability, predictability), and
point out by number which steps in generation_log count as 'creative contribution'. Anything you
cannot settle legally, mark as 'legal review needed' and hand it to a human. As a table: classification | basis (control) | basis (predictability) | contribution step | legal review.

Note that half the prompt is "forced reasoning" and a "legal escape hatch." Hand the AI a legal judgment wholesale and it will spit out a plausible-sounding verdict. The core of this prompt is the classification scheme (the guide's three categories), the reasoning requirement (the two axes), pointing at the contribution steps, and the escape hatch: "if you can't settle it, hand it to legal."

Step 3 — Claude's Output (Reconstructed from the Actual Session)

| Classification | Basis (controllability)              | Basis (predictability)            | Contribution step | Legal review |
|----------------|--------------------------------------|-----------------------------------|-------------------|--------------|
| Partial        | Rough sketch (step1) and 40% retouch | Picking 1 of 4 was predictable;   | 1,3,4             | ★            |
|                | (step3) directly control the expressive intent | the AI's 4 variations themselves were not |   |              |

Analysis:
- step1 (own rough sketch as input): falls under guide path ① 'creativity reflected by inputting one's own work'.
- step3 (40% retouch): falls under guide path ② 'creative modification/augmentation work'. The strongest contribution.
- step4 (hand-designed sealing-array sigil): an independent work composited in — clearly human creation.
- step2 (AI's 4 variations): pure output with no human creative contribution → not registrable on its own.

★ Legal review needed: there is no quantitative standard for whether a "40% retouch" reaches
  'substantial transformation'. Attorney review is recommended on whether the retouch ratio
  crosses the registrability line. (The AI cannot determine the legal sufficiency of the ratio.)

The most valuable part of the output is not the table but the ★ legal review line at the bottom. The AI reported its own limit — "I cannot determine whether a 40% retouch is legally sufficient" — and handed it to a human. A good prompt makes it possible for the AI to say it doesn't know. Not auto-reject — auto-report.

Step 4 — Verification and Rejection (the Human's Seat)

You don't accept this output as is. The director runs it against the rulebook once. The AI classified step4 (the sealing-array sigil) as an "independent work," but rereading the generation history, that sigil derived from the city lore that §6.2's city_hunting_generator produced. So step4 may not be pure creation but secondary work layered on top of an internal company asset. Since it is a company asset, ownership of the rights is clear — but the AI's phrase "independent work," written verbatim into the registration application, invites misunderstanding.

So I re-request.

The sealing-array sigil in step4 is secondary work derived from in-house city lore assets (not independent new creation).
Reclassify the nature of step4's contribution to reflect this fact.
Also propose, in one line, how the registration application should state that it is 'based on existing in-house assets'.

One round trip closes it. The AI reclassified step4 from "independent work" to "derivative work of in-house lore assets — rights to the source asset stay with the company; the transformative contribution is registrable," and that judgment went to legal review. The conclusion was confirmed: partial registration plus AI-generation disclosure. Done entirely by hand, legal would have to interrogate the generation history of every single asset; with an AI draft, a rulebook review, and one round trip, legal spends its time only on the borderline cases marked ★.

This full loop is this chapter's bar for Show. The sentence "AI copyright is murky" stays hollow until you have classified one output's generation history all the way through against the guide's criteria.


22.4.3 Decision Tree — Can This Output Be Used?

To avoid redoing the session's judgment from scratch every time, record the guide's criteria as a flowchart. When an asset comes in, you walk down this tree. Every branch point is a public criterion from §22.4.1.

flowchart TD
    A["AI output produced"] --> B{"Did a human input
and control the intent?
(rough sketch / own work / detailed direction)"} B -->|No, prompt only| C["Not-registrable output
→ exploration/concept reference only
no direct use as a final asset"] B -->|Yes| D{"Was the output modified/augmented
or selected/arranged?"} D -->|No| C D -->|Yes| E{"Is the model's training data
disclosed?
or an in-house fine-tune"} E -->|No or unclear| F["Hold for legal review
decide after infringement risk assessment"] E -->|Yes| G{"Derived from existing
in-house assets?"} G -->|Yes| H["Derivative work
source-asset rights stay in-house
+ register the transformative contribution"] G -->|No| I["Partial/full registration possible"] H --> J["Disclose AI generation
(AI Framework Act duty)
+ record in _economy_log"] I --> J F --> J classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; classDef fail fill:#fee2e2,stroke:#dc2626,color:#7f1d1d; class A ai; class B,D,E,F,G human; class J data; class H,I pass; class C fail;

The key point is that the tree's end (J) is the same on every path. Registrable or not, company asset or derivative work, the disclosure that AI was used and the generation-history log are recorded without exception. Disclosure is a duty separate from rights, and the log is the only basis for tracing responsibility when something goes wrong. The opening meeting stalled precisely because this log didn't exist — nobody could reconstruct who did what at which step.

The red path (C, not registrable) isn't simply thrown away either. "Pure AI output from a prompt alone" is perfectly usable as reference in the exploration and concept stage. You just don't put it into the game as a final asset. Shipping raw AI output unchanged is the single biggest opening for a post-release copyright incident.


22.4.4 The Copyright Rulebook as an Operating Log — design_intent_vs_automation_boundary

The tree (§22.4.3) is the flow of judgment; what makes that flow draw the same line every time is a single atom. Among the company's governance assets, design_intent_vs_automation_boundary is the spine of this entire chapter.

The atom's one-line definition is: "design intent belongs to people, automation to tools — make that boundary explicit per asset." It is not an abstract slogan. This atom is registered in the JIT hook (inject_memory.py), so whenever a prompt contains keywords like "copyright," "AI-generated," or "asset registration," it is automatically injected into the session. The hook's design principles directly support how this atom operates.

# inject_memory.py — always exit 0; even on failure, never block the user's flow (excerpt)
def main() -> None:
    ...
    # sort by score descending, then match — inject at most 3 atoms
    atoms_sorted = sorted(atoms, key=lambda a: a.get("score", 0), reverse=True)
    matches = []
    for atom in atoms_sorted:
        if len(matches) >= max_matches:   # prevent over-injection
            break
        try:
            if re.search(atom["regex"], prompt, re.IGNORECASE):
                matches.append(atom)
        except re.error:
            continue
    if not matches:
        emit_empty()   # no matches → empty response (normal)
        return

Two design decisions matter here for governance. First, the hook always exits 0 (stated in the script's docstring). Even if injecting the copyright rules fails, it never blocks the user's work. When a safety device holds work hostage, the team switches that device off within a quarter or two. Second, at most three atoms are injected. Push every governance rule into every session and the context blows up and nobody reads any of it. Only the highest-scoring rules surface.

This is the same philosophy as §6.2's lint, which did not auto-discard violations but only raised alerts to the writer gate. The machine picks the suspects; a human decides whether they live or die. The same holds in copyright: the atom automatically raises "did you check this asset's copyright?" — but the final judgment on registrability belongs to a human and to legal.


22.4.5 After Rights Come People — Closing Role Evolution with Agreement

The opening illustration's judgment did not end with copyright. It also meant that team member A's job had shifted from "drawing it" to "choosing among 4 AI variations and retouching 40%." The moment copyright demands "human creative contribution," the role definition of the person making that contribution changes with it. The two are the front and back of the same event.

The most common accident here is handling this change as an announcement. When a good tool sits unused six months later, it's usually not that the tool was bad — there was no agreement. The essence of adoption is roles shifting from mass production to selection, review, and reconstruction; if you don't state that explicitly and back it with training, team members hear it as "my seat is disappearing."

Role Before AI After AI (role evolution) Copyright meaning
Concept artist Draws everything by hand Inputs intent, selects, retouches The retouch is the 'creative contribution'
Balance designer Manual simulation Interprets sims, decides The decision log is the basis of accountability
Game designer Writes every spec Provides intent, reviews Intent input is controllability

The table says one thing: the "human contribution" that makes copyright registration possible is exactly what the person does after role evolution. When the control and prediction the guide demands disappear, the copyright disappears and so does the person's seat. So role evolution has to be explained not as a change that takes jobs away but as a change that keeps rights and responsibility in human hands — that's when agreement happens.

Agreement is not an endless meeting. You close it with a procedure: adoption proposal (director) → advance sharing with the whole team (purpose, affected roles, metrics, risks) → an agreement meeting (open floor, collect concerns) → 1:1s with members who need them → announce the adjusted plan → agree or hold. You don't need every member's consent to start, but you hear the concerns through a procedure, adjust, and then the director decides. Without the procedure, agreement restarts from zero every time, and that cost delays adoption.


22.4.6 Cost and ROI Are Part of Ethics — Honest Measurement with _roi_report.md

Narrow ethics down to jobs and agreement and you miss one axis. Honestly measuring and publishing the cost and effect of AI operations is itself governance. Say only "AI improved our efficiency" with no measurement, and team members will suspect the claim is a pretext for shrinking their seats.

The company's governance infrastructure has two real assets for this: the atom system's _economy_log/ (the token-and-time economics log) and _roi_report.md (the ROI report). The former has the machine record each session's tokens and time; the latter aggregates them on a cycle for humans to read. The point is that these logs track not "how much AI replaced people" but "where it freed people's time to go."

This book's rule for numbers is one of three. First, public standards are quoted as is (the guide's registration requirements, the AI Framework Act's disclosure duty). Second, author estimates are labeled as estimates. Third, only what's measurable is promised as a KPI. In copyright and ethics, what's measurable is not outcome metrics but procedural metrics.

Metric How measured Can it be promised?
Missing AI-generation disclosures grep asset meta for ai_generated_disclosure Measurable (target 0)
Generation-history log coverage Share of assets with a generation_log Measurable
AI assets shipped without legal review Release build vs. legal-cleared list Measurable (target 0)
"Revenue went up because of AI" Not measurable; not promised

The last row is the heart of honesty. AI adoption's revenue effect cannot be isolated as a single variable, so I do not claim causation. What can actually be counted, via _economy_log and the asset metadata, is "the share of AI-generated assets that passed disclosure, logging, and legal review." What governance promises is not outcomes but the integrity of the procedure.


22.4.7 User-Generated Content (UGC) and Data Protection

Settle asset rights, roles, and cost, and one area remains: the path where users put AI-made content into the game. Company-made assets close through internal procedure, but UGC pours in from outside your control.

The end of the §22.4.3 tree (disclosure and logging) applies here unchanged. User-uploaded costumes and guild emblems are required to carry an AI-generation disclosure, and moderation combines automated review with a human gate. Run only one of the two axes and the next quarter's incidents accumulate. And user data is never sent to an LLM carelessly: personal and payment data must not be transmitted at all, and behavioral logs go out only after anonymization (compliant with GDPR and Korea's Personal Information Protection Act).

Jurisdiction doesn't end with Korea. The moment you accept overseas users, the data regulations of each user's region attach as well. For EU users, GDPR sets separate requirements on cross-border transfer, consent, and the right to erasure; other service countries have their own privacy and data-localization rules. That's why the table's third row, "no personal or payment data to the LLM," is the safest default in any jurisdiction, and the strength of anonymization and pseudonymization when sending behavioral logs to an external model has to be checked separately for each region you serve. That said, this section is procedural design guidance, not legal advice. If a global launch or cross-border transfer is involved, get separate legal review in the relevant jurisdiction without fail.

UGC/data Policy Basis
User-uploaded costumes and emblems AI disclosure + automated review + human gate AI Framework Act disclosure duty
Character nicknames and posts Standard terms of service + user responsibility
Personal and payment data No LLM transmission Personal Information Protection Act
Game behavior logs Transmit after anonymization Anonymization/pseudonymization

The more UGC grows, the heavier the moderation burden gets. Automated review filters first; only the borderline cases reach a human. This is §22.4.4's atom philosophy (the machine picks the candidates, the human decides) carried over to the user-content level.


22.4.8 Common Failures

Pattern Why it fails Remedy
Using AI output unchanged as the final asset Zero human contribution → not registrable + infringement risk §22.4.3 tree; exploration/concept only
No generation-history log No way to reconstruct per-step responsibility after an incident Make generation_log mandatory in asset meta
No AI-generation disclosure Violates the AI Framework Act's transparency duty The disclosure step at the tree's end has no exceptions
Handling role evolution as an announcement Tool rejected six months after adoption Agreement procedure (§22.4.5)
Only shouting "AI raised efficiency" Team members suspect a threat to their seats Publish measurements via _economy_log and _roi_report
Uncritically using models with unclear training data Infringement risk assessment skipped Prefer disclosed models and in-house fine-tunes (tree branch E)

The fifth is the one most often missed. As the opening illustration's judgment showed, the "human contribution" that makes copyright registration possible is that person's new role. Measure only efficiency and not where that person's time was freed to go, and governance succeeds on the KPIs while the people leave.


Beyond Games. A meeting freezing over "we made this with AI — is the copyright ours?" happens not only with game illustrations but with AI-made reports, ad copy, and proposals anywhere. The standard the Korea Copyright Commission's 2025 guide nailed down is simple: registrability hinges on "human creative contribution (controllability and predictability)," and pure AI output from a prompt alone carries no rights. So in any department, the habit of leaving one generation-history sheet per AI output — which steps were human, which were AI — becomes a safety net. If a marketer, for example, takes an AI copy draft and personally revises and reconstructs it, recording that contribution becomes the basis for claiming rights; and whether or not it gets registered, the disclosure that AI was used (a duty under the 2026 AI Framework Act) is recorded without exception. After rights come people: the role change for the employee making that contribution should be closed by agreement, not by announcement.

22.4.9 Try It Yourself — One Step You Can Take Today

If you're solo, just this much: You don't need a legal team. Pick one image or text output you made with AI and write out a generation_log in the §22.4.2 format by hand (split it into steps: which were human, which were AI). Then walk the §22.4.3 tree and judge for yourself whether it's registrable — the Korea Copyright Commission guide's "control and prediction" criteria will sink in as a concrete bundle of judgments. Even on a personal or hobby project, it's worth leaving the one-line AI-generation disclosure (ai_generated: true).

If you're on a team, start with this one step. Make the two slots generation_log and ai_generated_disclosure mandatory in every AI asset's metadata (a one-line grep catches omissions), and pin the §22.4.3 decision tree to your wiki as a single page. Automating registrability judgments or running _economy_log comes later. With just the generation-history log and one tree page, you can prevent the kind of meeting that froze at the opening.


22.4.10 Wrapping Up Part 22

Part 22 covered the four axes of governance.

Chapter Core
22.1 Prompt engineering — forcing format, reasoning, and escape hatches
22.2 Hallucinations and safety — review gates, report-style verification
22.3 Cost management — caching, caps, _economy_log
22.4 Copyright and ethics — registration requirements, disclosure, role agreement

One sentence runs through all four chapters: governance is not a device that blocks AI; it is the procedure that leaves human intent and responsibility on the output. The design_intent_vs_automation_boundary atom is that procedure's name. The "control and prediction" that copyright registration demands, the "roles and agreement" that ethics demands, and the "honest measurement" that cost demands all point to one place — whatever the AI does, the last seat of decision and responsibility belongs to a human.


Key Takeaways

Next Chapter Preview



Sources - Korea Copyright Commission and Ministry of Culture, Sports and Tourism, "Guide to Copyright Registration for Works Made Using Generative AI" (2025) — https://www.copyright.or.kr/information-materials/publication/research-report/view.do?brdctsno=54253 - Korea Copyright Commission and Ministry of Culture, Sports and Tourism, "Generative AI Copyright Guide" (Dec. 2023) — https://www.copyright.or.kr/information-materials/publication/research-report/view.do?brdctsno=52591 - The AI Framework Act (Korea's basic law on AI): transparency and disclosure duties for generative AI outputs (effective 2026) — https://www.shinkim.com/kor/media/newsletter/3142

Part 23 · Extension

Part 23 · Chapter 1. The Wrapper, Cascade, and Junction Patterns

Don't add more tools — build tools for your tools. This is the story of a two-tier structure that hides 48 implementations behind 12 global entry points, and the automation that keeps the two consistent without human hands.


One evening, while running my monthly retrospective, I was counting my slash commands and my hand stopped. Forty. I clearly started with seven or eight half a year earlier, but I built a meeting-notes tool here, attached a data-validation tool there, added a game design document (GDD) generator — one or two a week — and somewhere along the way the list had grown to forty. And nearly half of them I hadn't called even once in the past month.

The problem was that unused tools don't just sit there quietly. Every time a session started, all forty slash command specs were loaded. They ate into the token budget, similarly named commands (skill-design, skill-design-new, skill-design-template) blurred together, and recalling the tool I actually needed took time. The tools weren't helping the work anymore — managing the tools was becoming the work.

This chapter covers how I cut those forty down to twelve global commands without throwing away a single implementation. Three patterns carry the work: Wrapper, which creates lightweight entry points; Cascade, which bundles multiple tools behind one entrance; and Junction, which physically links entry points to their implementations. Plus sync_skills.py, which guards the consistency of all three so a human doesn't have to.


23.1.1 The Quantitative Signal Found in the Retrospective

Everyone has the impression of having too many tools. But an impression alone can't decide what to cut. What made the decision possible was the tool-economy measurement in the monthly retrospective.

This project runs retrospectives as a self-improvement mechanism. Daily retrospectives accumulate into weeklies, weeklies merge into monthlies, and along the way the monthly retrospective back-calculates "which tools did I use, and how often, over the past month" from the SVN commit log. The score used for this measurement is skill_audit_score. It tracks how often each slash command appears in actual work artifacts through commit history and assigns a usage frequency.

The distribution that month's measurement revealed looked like this. (The usage percentages are measured from the SVN commit log; they are per-tool shares of appearances, not absolute call counts.)

40 slash commands — usage frequency distribution Top 12 commands 92% of usage 10 mid-usage commands — about 8% 18 used less than once a month (45% of the total) — nearly 0% Source: monthly retrospective skill_audit_score, back-calculated from the SVN commit log / percentages are measured shares of appearances

The top twelve accounted for 92% of all usage, and eighteen commands — 45% of the total — weren't used even once a month. Half the answer was already decided: expose only the twelve frequently used commands globally and clean up the rest.

The catch was that "clean up" didn't mean "delete." Even the twenty-eight unused commands were needed once or twice a quarter — when writing a half-year report, building a new data schema, or running a specific validation. If the tool isn't there at that moment, the work stops on the spot. So the real question was this: how do I show only twelve while keeping all twenty-eight alive?

A desk metaphor runs through this entire chapter. Nobody lays out forty pens on their desk and uses them all every day. You keep the twelve you use most on the desk and put the rest in a drawer. Inside the drawer, pens of the same kind go into one cup. Wrapper is the lightweight entry point you keep on the desk, Junction is the passage connecting the drawer to the desk, and Cascade is the bundle of pens gathered in one cup.


23.1.2 The Wrapper Pattern — Lightweight Entry Point, Heavy Implementation

A Wrapper is a thin shell around a slash command. Only the entry point lives in global; the actual logic lives in the implementation under workspace. The global directory holds a 50-line guide; the implementation holds a 500-line build-out.

flowchart LR
    subgraph G["Global ~/.claude/skills/  (on the desk)"]
        W1["proj-meeting
Wrapper · 50 lines"] W2["proj-gdd
Wrapper · 50 lines"] end subgraph B["workspace/skills/ (in the drawer)"] M1["proj-meeting/
SKILL.md + extraction/classification .py
about 500 lines"] M2["proj-gdd/
SKILL.md + generator
about 500 lines"] end W1 -->|calls| M1 W2 -->|calls| M2 classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; class W1,W2,M1,M2 code;

This separation pays off in five ways. Only 50 lines load globally at session start, saving tokens; the implementation can be edited daily without touching the global slot; the implementation can live anywhere — SVN, Git, wherever; sharing is easy because the implementation sits in a shared team folder while only the Wrapper lives in your personal global directory; and a unified Wrapper format keeps the user experience consistent.

The standard Wrapper format looks like this. Every Wrapper shares this skeleton.

---
name: proj-meeting
description: Meeting-notes analysis and decision extraction (implementation: workspace/skills/proj-meeting/)
---

# /proj-meeting — Wrapper

Implementation location: workspace/skills/proj-meeting/SKILL.md

## Behavior
This Wrapper calls the implementation's entry script. Detailed logic is defined in the implementation.
When the implementation changes, only this Wrapper's description needs updating (automatic sync recommended).

The point is that there's nothing but a one-line description and a pointer to the implementation. The moment logic creeps in, the Wrapper gets heavy and synchronization with the implementation starts to break. So I enforce a rule: a Wrapper stays under 100 lines.

This project caps its global slash command slots at twelve. All frequently used tools must fit within those twelve, and the monthly retrospective owns the selection criteria: used five or more times a month, balanced across domains (no single domain exceeds six tools), and consistent entry (a unified naming convention). When the count exceeds twelve, the least used one is retired or merged into another command.

Twelve is not an absolute number. What matters is that a number is fixed at all. A small team (up to about 10 people) might do fine with ten; a team spanning many domains might be right at fifteen. Only a fixed ceiling keeps the cognitive load at a constant level.


23.1.3 The Junction Pattern — Physically Linking the Implementation and the Entry Point

If Wrapper is the rule — "only lightweight entry points go global" — Junction is the means of implementing that rule at the operating-system level. A Junction is a directory symbolic link: an alias provided by the OS.

flowchart LR
    U["~/.claude/skills/proj-meeting
(Junction — alias)"] R["workspace/skills/proj-meeting/
(implementation — the single real copy)"] U -. "actually points to" .-> R classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; class R code; class U data;

When the user looks into the global location, the implementation appears to be sitting right there. But the actual files exist in exactly one copy, at the implementation's location. The global side is merely a signpost pointing there.

The gains from this structure are clear. Edit the implementation and the change shows up globally at once (there is no copy step). Files exist in one copy, saving disk space, and the global side holds only Junctions, so there are no Git conflicts (the implementation is managed separately in SVN or Git). Move the implementation, re-point the Junction, and the user notices nothing.

How you create one differs by OS. On Windows, mklink /J <link> <target> creates a directory junction, and no administrator rights are required. Linux and macOS use ln -s <target> <link>, and WSL uses the Linux command as is. The sync_skills.py script covered below handles this platform difference automatically, so the operator never has to memorize per-OS commands.

Run on copies instead of Junctions, and a synchronization accident happens the moment the implementation and the global copy diverge: you fix a bug in the implementation, but the global copy is an old version and keeps the old behavior. A Junction eliminates the very possibility of that accident. A signpost can't become a second copy, and the substance is always one.


23.1.4 sync_skills.py — Keeping Consistency So a Human Doesn't Have To

Manage Wrappers and Junctions by hand and you eventually drift back to forty. People postpone cleanup, forget policies, and make exceptions. So consistency maintenance is automated. That tool is sync_skills.py.

A hook triggers this script every time a session starts. What the script does flows like this.

flowchart TD
    H["Session start (hook trigger)"] --> S["Scan ~/.claude/skills/"]
    S --> C{"Matches the
twelve-Wrapper policy?"} C -->|"leftover slots found"| X["--cleanup:
remove off-policy Wrappers"] C -->|"implementation move detected"| J["Auto-recreate the Junction"] C -->|"matches"| OK["Pass"] X --> OK J --> OK OK --> R["Global 12-slot consistency guaranteed"] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; class H,S,C,X,J code; class OK,R pass;

It has three core functions. First, it scans the global directory and checks compliance with the twelve-Wrapper policy. Second, with the --cleanup flag it removes leftover Wrappers that aren't in the policy. If a tool someone added temporarily lingers in a slot, it gets cleaned up at the next session start, so the slots never balloon again. Third, if an implementation's location has changed, it re-creates the Junction automatically — detecting the OS and calling mklink /J on Windows, ln -s everywhere else. What matters is that all three functions are designed to be idempotent. Since the tool runs automatically at every session start, running it any number of times on the same state must produce the same result as running it once. Wrappers that already match the policy are left alone, Junctions that are already set correctly are not re-created, and when there are no leftover slots to clean, nothing is deleted. Without idempotence, the same cleanup would pile up session after session, re-creating healthy Junctions or touching implementations it shouldn't — and for a tool that runs unattended every session, that leads straight to synchronization accidents. So sync_skills.py holds one invariant: touch only what changed; if nothing changed, touch nothing.

The effect of --cleanup translates directly into token-budget protection. Pin the slash specs loaded globally each session at twelve, and even as the implementations grow to forty-eight, the session-start cost stays constant. Because no human manages it by hand, the policy never drifts.

This automatic consistency is the safety pin of the two-tier structure. Wrapper and Junction build the structure, and sync_skills.py keeps that structure standing as time passes.


23.1.5 The Two-Tier Structure — 12 Global Wrappers → 48 Workspace Implementations

Put the three patterns and automatic consistency together and the following two tiers are complete. On the upper tier sit the twelve entry points the user memorizes; on the lower tier sit the forty-eight implementations.

flowchart TD
    subgraph L1["Tier 1 — 12 global Wrappers (all the user memorizes)"]
        direction LR
        w1["#1"] -.- w12["#12"]
    end
    subgraph L2["Tier 2 — 48 workspace implementations (the hidden substance)"]
        direction LR
        b1["Implementation 1"] --- bN["Implementation 48"]
    end
    L1 -->|"linked via Junctions"| L2
    note["sync_skills.py --cleanup:
aligns tier 1 to 12 every session"] note -.-> L1 classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; class w1,w12,b1,bN,note code;

The user remembers only the twelve global commands. Even with forty-eight implementations hidden behind them, the cognitive load stays at twelve. Wrapper keeps the entry points light, Junction links entry points to implementations, and sync_skills.py guards the consistency of those twelve every session.

As a ratio, entry points to implementations is 1 to 4 (12 to 48). Adding tools doesn't add to what the user must memorize. Grow the implementations to sixty or eighty, and tier 1 is still twelve. This is the actual implementation of the sentence "don't add more tools — build tools for your tools." What grows is tier 2 (the implementations); tier 1 (the entry points), the part the user faces, stays constant.


23.1.6 The Cascade Pattern — Chained Calls Behind One Entrance

If the two-tier structure is the pattern that "reduces many tools to few entry points," Cascade is the pattern that "bundles tools often used together into a single call." One slash command calls multiple sub-tools in sequence and produces one combined report.

This project's flagship Cascade is check. It consolidated into one the four tools that used to inspect the integrity of design data every morning.

flowchart TD
    E["/check  (Wrapper · Cascade entrance)"] --> S1["doc-audit
Markdown consistency"] S1 --> S2["data-qa
data sheet validation"] S2 --> S3["integrity
foreign key consistency"] S3 --> S4["link-check
Wikilink integrity"] S4 --> R["Combined report
(failures in detail, passes summarized)"] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; class E,S1,S2,S3,S4 code; class R data;

I used to call the four tools separately every morning: one document check, one data check, one foreign key check, one link check. Each work cycle cost three or four manual calls. check bundles the four into one command — call it once and the four steps run in order, with the results merged into one.

Cascade design follows principles. Each step must be callable on its own (you have to be able to call just data-qa by itself). Whether to stop or continue on failure is set per step: validation work keeps running the remaining steps after a failure to see the whole picture, while modification work stops immediately when a step fails. Results accumulate and become the next step's input, and the combined report uses the same format across every Cascade.

Here is the actual definition of check. (It consolidates four kinds of validation into one.)

cascade:
  - step: doc-audit
    purpose: Markdown consistency (YAML frontmatter, links, atom references)
    fail_action: continue
  - step: data-qa
    purpose: Excel data sheet validation (schema, ranges, required columns)
    fail_action: continue
  - step: integrity
    purpose: foreign key consistency (cross-sheet references)
    fail_action: continue
  - step: link-check
    purpose: Wikilink and external link integrity
    fail_action: continue

report:
  format: markdown
  include_pass: false   # passes summarized only; failures in detail
  group_by: severity

fail_action: continue is set on all four steps because this is a validation Cascade. Even if one check fails, the other three still run, so I see that day's full defect list in one pass. The report folds passing items into a summary and expands only the failures, gathering the morning's attention on what needs to be seen.

Cascade has its own trap. Add steps without limit and complexity explodes. So, like the 12-slot policy, Cascades get a step ceiling. Past roughly five to seven steps, split it in two or move some steps into a separate Cascade.


23.1.7 Tool Curation — The MECE Wrapper Policy

Once the two-tier structure and Cascade settle in, growing the implementations becomes easy — just add one to workspace without touching the global slots. And right there a new trap appears: when adding is easy, similar tools pile up as duplicates.

So one policy is enforced when growing the implementations: the MECE Wrapper policy. When a new tool is about to be added, the judgment forks two ways. If its territory overlaps an existing tool, don't create a new one — augment the existing tool. Only when the territory is clearly different is a new one created. It means keeping the implementation list Mutually Exclusive (no duplicates) and Collectively Exhaustive (no gaps).

flowchart TD
    N["Need a new tool?"] --> Q{"Overlaps an existing
implementation's territory?"} Q -->|"overlaps"| A["No new tool →
augment the existing one"] Q -->|"clearly different"| B["New implementation allowed"] A --> M["MECE preserved: no duplicates"] B --> M classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; class Q human; class M pass;

The retrospective backs up this judgment. Because skill_audit_score measures per-tool usage frequency from the SVN log, signals like "this tool does nearly the same thing as that one, and neither gets used much" get caught. Then the two are merged into one, or the less used one is retired from the implementations. In curation, removal matters as much as addition.

Without the MECE policy, the "freedom to add implementations" that the two-tier structure created turns into poison instead. Even if implementations never eat into the tier-1 slots, once the implementations themselves bloat with duplicates, you start getting confused again about which implementation to use. The policy blocks that bloat.


23.1.8 Why the Retrospective Ignited All of This

Let's go back to the beginning. Not one of these — Wrapper, Junction, Cascade, the MECE policy — was designed up front at a desk. Every one of them arose as an answer to a problem found in a retrospective.

When the monthly retrospective's skill_audit_score exposed, in numbers, the forty slots and the 92% usage concentration, "cap it at twelve" was decided. That decision dragged along the question "then how do we keep the twenty-eight alive," and the answer was Wrapper and Junction. The next retrospective surfaced the finding "calling four similar validation tools separately every morning is a chore," and the answer was the check Cascade. Yet another retrospective caught the signal "now that adding implementations is easy, duplicates are piling up," and the answer was the MECE policy.

Without retrospectives, these patterns would not have been built. Or if built, they would have become over-engineering detached from real problems. Finding the problem through measurement first, then introducing the pattern — that order is what makes tools actually get used. Part 21's message that the retrospective is the starting point of self-improvement takes concrete form here, at the level of tools.


23.1.9 Operating Record — Six Months Accumulated

Here are six months of measurements compared before and after adoption. The focus is the change in operating burden, not absolute call counts.

Item Before After (Wrapper+Cascade+Junction)
Global slot count 40 (ballooning) 12 (policy-enforced)
Implementation count scattered, many duplicates 48 (MECE-curated)
Global slot share at session start large (40 specs loaded) small (only 12 specs loaded)
Global propagation after editing an implementation manual copy step required immediate (Junction, no copying)
Calls to similar validation tools 3–4 manual calls per task 1 (check Cascade)

The first month's measurements were uneven. There were a couple of synchronization accidents before the Wrapper format settled, and the slots hovered between fifteen and eighteen before the twelve-cap policy was enforced. Stabilization came in month two. Once sync_skills.py --cleanup started aligning the slots every session, the slot count never ballooned again.

The directional wording in the table — "share," "required," "immediate" — is deliberate. Token costs and time differ by environment, so I recorded only the direction of change. What is certain is that tier 1 went from forty to a fixed twelve, and the manual copy step disappeared from implementation propagation.


23.1.10 Common Mistakes and How to Avoid Them

Here, gathered in one place, are the traps flagged in the preceding sections. All five point to the same lesson: keeping a structure alive over time is harder than building it.


Try It Yourself (setup → prompt → verify)

setup. Create one implementation directory under workspace (e.g., workspace/skills/proj-meeting/). Put SKILL.md and the actual scripts inside it. In the global ~/.claude/skills/, place only a 50-line Wrapper.

prompt. Ask Claude the following.

Scan the list of slash commands in ~/.claude/skills/.
Classify each command as (a) a lightweight Wrapper holding
only a pointer to its implementation, or (b) a heavy command
with logic inside; for each (b), propose a change that moves
the implementation into workspace and leaves only a 50-line Wrapper globally.
Then print the command that creates the Junction pointing to the
implementation, matched to the OS (mklink /J on Windows, ln -s otherwise).

verify. Check three things. First, is each item in the global directory under 100 lines? Second, when you open a global item, do you see only one line with the implementation's location plus a description? Third, edit one line in the implementation and call it from global — does the edit show up immediately? (If the Junction is set correctly, it propagates with no copying.)

Solo Scale-Down

If you're a solo operator with no team and no SVN, scale it down like this. You don't need a workspace — one personal Git repository is enough. Keep the implementations in that repository and only Wrappers in global. Even without a measurement tool like skill_audit_score, simply writing down by hand at month's end "the commands I actually called this month" reveals the concentration. Keep only the top five to seven in global and move the rest down to implementations. For Cascade, bundling into one command once two or more tools are routinely called together is enough. If an automatic consistency script feels like too much, replace it with the habit of glancing over the global directory once at the start of each new session. At a smaller scale, habit does the same job automation does.


Key Takeaways

Next Chapter Preview

Part 23 · Chapter 2. Adopting Hermes Agent

11:47 p.m. I saved the data sheets one last time and closed the laptop. At 9:10 the next morning, while the coffee brewed, I opened the team messenger and found a report sitting at the top of the channel. A markdown file that had cross-checked the three balance data sheets updated overnight against their foreign keys, with two broken references flagged in red. I did not write it. It was made while I slept.

This chapter is the record of the tool that produced that report — putting Hermes Agent on my personal PC and layering it on top of the Wrapper, Cascade, and Junction operations covered in §23.1. When I first brought it in, Hermes was Linux-based, so using it on Windows meant going through WSL2; in 2026 a native Windows build arrived and that detour disappeared. The conclusion first: the agent did not push Claude Code aside. It took the seat next to it.


23.2.1 Two Tools at the Same Desk

Everything up through §23.1 was centered on Claude Code. I typed one sentence, the tool responded once, I reviewed the response, then typed the next sentence. This short cycle is unbeatable for precision work. If fixing a single balance number is a task that needs confirmation at every step, then a human stepping in every time is exactly right.

The problem was long-running work. A request like "read all 30 of last month's meeting notes and extract just the decisions as atom candidates" takes 30 round trips if handled inside a conversation flow. For those 30 round trips, I can do nothing else. For this kind of work, the strength of a tool whose inputs and outputs sit close together becomes a weakness.

The agent fills the opposite seat. Throw it only a goal — "pull the decisions out of 30 meeting notes as atom candidates and write a report" — and it picks its own tools, walks the intermediate steps on its own, and brings back only the result when it is done. The cycle is long and autonomous. The trade-off that comes with it: a human cannot watch every step.

Work Cycles of the Two Tools Claude Code Precise · short · verified each step in→out in→out in→out in→out ↑ a human reviews at every arrow Hermes Agent Long-running · autonomous · checkpoints only 1 goal in → (autonomous run: tool choice · loops · checks) → 1 result out ● checkpoints (where a human can review) — not every step Precise decisions on the top lane, long repetitive work on the bottom. Same desk.

An office analogy makes it easier. Claude Code is the desk mate who reads every sentence alongside me; the agent is the assistant who volunteers for the night shift and leaves a report on my desk before I arrive. Neither one fires the other. The two share the same desk.


23.2.2 Why Bring In Yet Another Tool

In §23.1 I built an operation that bundles the global slash command slots into 12 and hides the 48 actual commands behind them with Junction. The conclusion there was to build tools for the tools instead of adding more tools. Bringing in yet another new tool here sounds like a contradiction of that conclusion.

It is not. The 12-slot policy of §23.1 dealt with the cognitive load of "tools a human invokes directly." The seat Hermes wants to fill is the hours when no human is invoking anything — the hours I am asleep, in meetings, or with my hands tied up elsewhere. It does not compete with the 12 slots; it fills the hours the 12 slots cannot reach.

The basis for the adoption decision was one retrospective measurement. Running a month of data through skill_audit_score, which back-calculates global tool usage frequency from SVN commit logs, showed that most of the top tools were the kind used "while a human is awake, briefly, and often." Meanwhile, the tasks with low usage frequency but long runtimes once started — batch-classifying meeting notes, nightly data sheet integrity, build capture analysis — kept getting postponed with "I'll do it tomorrow morning." The reason for the postponement was clear. They eat up long stretches of waking hours.

This postponed group of tasks is the agent's exact target.


23.2.3 Installation — The Native Windows Build

When I first set up Hermes, it was Linux-based, so using it on a personal Windows PC meant installing WSL2 (Windows Subsystem for Linux 2) first and seating Hermes inside it. There is now a native Windows build, so that detour is no longer needed. Installation works like any ordinary Windows application — download the installer, run it, and set the initial workspace path and permission allowlist on first launch.

If you already use WSL2 or prefer a Linux environment, that build is still fully supported. But if you are starting fresh, native is simpler. The exact installer and version change quickly as the tool evolves, so follow the official docs.

One trap remains regardless of where you install. The Hermes workspace must live on a fast local disk. Wire a network drive or an SVN working folder directly in as the workspace, and a nightly integrity check that should take a few minutes stretches into tens of minutes. The standard practice is to keep data sheets outside the workspace and copy them in only when a task starts. If you use WSL2, for the same reason, keep the workspace inside the Linux filesystem and do not hop across Windows paths like /mnt/c.


23.2.4 Installing Hermes and the First Connection — A Worked Transcript

This is where the hands-on part begins. What you have the tool do after installation matters more than the installation itself, so I follow one first task all the way through — full prompt, raw output, human verification, re-request. For this task I picked the simplest of the postponed group from §23.2.2: the nightly data sheet integrity check.

Note: some of the commands below are illustrative, shown to convey Hermes's surface shape. Installer URLs and subcommands change between versions — check the official docs. The structure of the workflow (goal → autonomous run → verification → re-request) holds even as the tool changes.

On native Windows, you fetch the official install.ps1 in PowerShell and run it. But before executing the one-line iex (irm ...) one-liner as is, the safer path is to download the script (about 2,800 lines) once and skim it for dangerous patterns with your own eyes — that is the minimum procedure for trusting a source — then split off the key setup: install just the body first with -SkipSetup and run hermes setup separately. On WSL2 or Linux, follow the corresponding installation section of the official docs.

# Native Windows — official install.ps1 (download and review first, then run)
irm https://hermes-agent.nousresearch.com/install.ps1 -OutFile install.ps1
# (after reviewing the contents of install.ps1)
.\install.ps1 -SkipSetup
# Pulls in Python 3.11 · Node · Git · Playwright · bundled skills along the way
# Install location: %LOCALAPPDATA%\hermes\  (registers the hermes command on PATH — picked up from a new terminal)
# When it finishes: hermes setup

The installer brings the dependencies (Python 3.11 · Node 22 · Git) along, installs the body into %LOCALAPPDATA%\hermes\, and registers the hermes command on PATH (recognized from a new terminal onward). Operational data — settings, logs, schedules (cron), checkpoints — also stays under the same %LOCALAPPDATA%\hermes\ and survives reinstalls (here is the trap — ~/.hermes\ contains only helper scripts, which makes it easy to get confused. The actual config.yaml and logs\ all live on the %LOCALAPPDATA%\hermes\ side). Run hermes setup as the first launch and it asks for your model API key and sets the initial workspace path and permission allowlist.

hermes --version
hermes setup

Now I hand it the first task. The goal you throw at an agent is one level more abstract than a Claude Code prompt. Less "do this, like this" and closer to "have this result ready." The exact goal I entered was the following.

[Full Prompt]

Goal: nightly data sheet integrity check.

Read the three xlsx files item_master·drop_table·npc_shop in ~/hermes-workspace/sheets/,
and check the foreign keys: do the item_id values referenced by drop_table·npc_shop
actually exist in item_master? Find every broken reference (rows pointing to an
item_id that doesn't exist) and save the result as a date-named markdown file in
reports/. (Intent: passes as counts only; failures as a table of sheet name · row
number · item_id)

Constraints:
- sheets/ is read-only. Never modify it.
- If you can't read a sheet, or a column name differs from what's expected, stop and
  report what was different. No guessing.

The last two lines matter. Agents are autonomous, and when they hit a wall they tend to want to fill the gap "on their own." If a column name differs, they pick a similar-looking one and push on. So I explicitly write in: "if you get stuck, stop and report — do not guess." This is the same lesson, in the same grain, that I learned building the data sheet verification Cascade in §23.1.

[Hermes Raw Output — First Pass]

Carried over verbatim, no summarizing. The first attempt did not end cleanly.

[hermes] goal accepted: nightly data sheet integrity check
[hermes] step 1/4  reading sheets...
         item_master.xlsx   loaded (4,210 rows)
         drop_table.xlsx    loaded (1,880 rows)
         npc_shop.xlsx      ERROR

[hermes] npc_shop.xlsx: expected column 'item_id', found columns:
         ['shop_id', 'sell_item', 'price', 'stock']
[hermes] this conflicts with the constraint: "if a column name differs from what's
         expected, stop and report".
[hermes] HALTING. reporting instead of guessing.

[hermes] partial result:
  - item_master ↔ drop_table check is feasible (item_id column present on both sides)
  - npc_shop: the item_id column is presumably 'sell_item', but unconfirmed — not run
[hermes] awaiting clarification on npc_shop column mapping.

This is the point where first-time agent users are most surprised. The tool stopped. And it stated exactly why it stopped. The item reference column in npc_shop.xlsx was not the item_id I expected but sell_item. Had I not written the constraint in, the agent would have decided on its own to treat sell_item as item_id and pushed on, and all I would have received in the morning is a report saying "all clear." The halt is the safety.

[Human Verification and Rejection]

Looking at the output, I made two judgments.

First, the part reporting that the item_master ↔ drop_table check was already feasible is trustworthy. Second, the sell_item guess for npc_shop is correct — it really is that sheet's sale-item column name. But the call not to proceed on a "guess" was the right one. So I did not accept the first output as is (rejection) and submitted a re-request with the column mapping made explicit.

[Re-Request Prompt]

Confirmed. Here is the explicit column mapping:
- The item reference column in npc_shop.xlsx is 'sell_item'. Treat it as item_id and run the check.
- All other constraints unchanged. Do not guess any additional columns beyond 'sell_item'.
Run again.

[Hermes Raw Output — Second Pass]

[hermes] resuming with column map: npc_shop.sell_item -> item_id
[hermes] step 2/4  foreign-key check...
         drop_table → item_master   : 1,880 refs, 2 broken
         npc_shop   → item_master   : 642 refs, 0 broken
[hermes] step 3/4  collecting broken refs...
[hermes] step 4/4  writing report...
         ~/hermes-workspace/reports/2026-05-25_sheet_integrity.md  written
[hermes] done. 2 broken references found. summary:
  | sheet      | row  | bad item_id |
  |------------|------|-------------|
  | drop_table | 417  | ITM_90412   |
  | drop_table | 1356 | ITM_90727   |

Two broken references were caught. Rows 417 and 1356 of drop_table were pointing at an item_id that did not exist in item_master. The red table I saw that morning was exactly this.

Nearly the whole message of this chapter shows up in this one round trip. The agent was autonomous yet stopped at the constraint, and once a human filled the spot where it stopped, it went all the way. Autonomy and control do not collide — they interlock. And if you schedule this entire cycle to run once more while you sleep, that becomes the nightly automation of §23.2.5.


23.2.5 Three Seats Added to the Game Design Workflow

Once the first task sits comfortably in your hands, you move the postponed tasks to the night shift one by one. I actually added three seats. What the three have in common is plain — every one of them turns hours when no human needs to be awake into working hours.

flowchart TD
    A["Nightly trigger
(daily 23:00, cron)"] --> B{Hermes Agent} B --> C1["[Seat 1] Data sheets
nightly integrity check"] B --> C2["[Seat 2] Long-run simulation
100 hours of virtual play"] B --> C3["[Seat 3] Build captures
automated analysis pipeline"] C1 --> D1["Foreign-key diff
broken-reference table"] C2 --> D2["Avg. boss-kill time · resource spend
combo distribution"] C3 --> D3["Spec vs. measured diff
per-frame extraction"] D1 --> R["[Combined] Markdown report
~/hermes-workspace/reports/"] D2 --> R D3 --> R R --> S["9:00 a.m.
auto-posted to team messenger channel"] S --> H["Game designer: reviews results only
(analysis finished while asleep)"] style A fill:#fff3e0,stroke:#e65100 style B fill:#e3f2fd,stroke:#1565c0 style R fill:#e8f5e9,stroke:#2e7d32 style H fill:#fce4ec,stroke:#c2185b

Seat 1 — nightly data sheet integrity. The task we followed end to end in 2.4, scheduled for 23:00 every night. Whoever touched whichever sheet overnight, by morning the broken foreign keys are sitting there as a table. On the surface this resembles what the /check Cascade of §23.1 did (the four-in-one bundle of doc-audit → data-qa → integrity → link-check), but there is one decisive difference. /check only runs when I am awake to invoke it. The nightly agent runs without me. The two do not compete — the daytime Cascade is immediate verification, the nighttime agent is unattended verification, and the roles split there.

Seat 2 — long-run simulation. The combat simulation from §4.4, stretched deep along the time axis. Run 100 hours of virtual play and measure average boss-kill time, resource consumption curves, and combo distribution. This fundamentally does not fit the Claude Code conversation flow — one run takes hours, and you cannot hold the chat window hostage for that long. The agent runs it in the background and brings back only the curve charts and summary numbers when it finishes.

Seat 3 — automated build capture analysis. When a build video captured by QA lands in the folder, the agent extracts data frame by frame and produces a diff between the spec numbers and the measured numbers. The designer does not need to watch the video end to end — just diff lines like "spec says damage 120, build measures 108." The entire tedious part of the analysis belongs to the agent.

In all three seats, the time spent by the person reading the results does not shrink. What shrinks is the human time spent on analysis. The judgment still belongs to the human.


23.2.6 The Price of Autonomy — Five Safeguards

The agent's autonomy is, in equal measure, risk. A tool that reads files and runs commands without a human confirming each step means that when things go wrong, the human is not there. It is no accident that "no guessing" was written explicitly into §23.2.4. The five safeguards are not optional — they are a bundle that must be switched on together on day one.

Safeguard What It Does (Actual Hermes Config Keys) What Happens Without It
Permission allowlist Destructive commands go through human approval (approvals.mode: manual), only allowlisted commands run free (command_allowlist), secrets are redacted from logs (security.redact_secrets) The agent autonomously modifies the original data sheets
Checkpoints Snapshots before file operations so you can roll back (checkpoints.enabled, restore with /rollback) A wrong assumption rolls all the way through and contaminates the entire result
Automatic logging Writes gateway, agent, and error logs to %LOCALAPPDATA%\hermes\logs\ After an incident, no way to trace why it happened
Cost limits Per-task turn cap (agent.max_turns), terminal timeout (terminal.timeout), infinite-loop auto-detection (tool_loop_guardrails), automatic context compression (compression) A task stuck in an infinite loop inflates the API bill
Disposability Stop anytime (/stop), pause/delete schedules (cron pause), sub-task timeouts (delegation.child_timeout_seconds), auto-archive of unused skills (curator) A nightly task that starts running wrong cannot be stopped

These five are not independent devices; they work as one bundle. Lock down permissions but skip the cost limit, and an infinite loop spins inside the permitted zone while the bill grows. Turn on logs but have no disposal path, and you watch the incident happen without being able to stop it. Drop any single one and the incident probability of unattended nightly operation jumps.

Actually turning the tool on, I found several places where it implements these five concepts one notch more finely than the book sketched them. On the permission side there is an extra layer, a dedicated policy engine (security.tirith_enabled), that filters commands by rule. On the cost side, infinite-loop detection is not a single cap but separate thresholds for signals like "same failure repeating" and "repetition with no progress." And unattended nightly scheduling (cron) has its own switch (approvals.cron_mode: deny): when a destructive command comes up during hours when no human is present, it is denied outright instead of waiting for approval — effectively the book's "permissions + checkpoints" folded into one setting. On the disposal side, curator is where §21's "retire the tools you don't use" ships as an actual feature. Keep the skeleton of the five-piece bundle as is, and where the tool is more refined, just turn those keys on.

Entering this bundle into config.yaml looks roughly like this.

# %LOCALAPPDATA%\hermes\config.yaml (excerpt)
approvals:
  mode: manual              # ① Permissions — destructive commands need human approval
  command_allowlist:        #    list only the commands allowed without approval
    - "python *"
    - "rg *"
  cron_mode: deny           #    unattended nightly cron auto-denies destructive commands
security:
  redact_secrets: true      #    redact secrets from logs
  tirith_enabled: true      #    one more layer: policy engine (rule-based command filter)
checkpoints:
  enabled: true             # ② Checkpoints — snapshot before file ops (/rollback to restore)
  max_snapshots: 20
  retention: 7d
logs:
  path: "%LOCALAPPDATA%\\hermes\\logs"   # ③ Logs — gateway/agent/errors
agent:
  max_turns: 60             # ④ Cost — turn cap per task
terminal:
  timeout: 180              #    terminal command timeout (seconds)
tool_loop_guardrails:       #    infinite-loop auto-detection (same failure · no progress)
  enabled: true
compression:
  enabled: true             #    automatic context compression (token savings)
delegation:
  child_timeout_seconds: 600  # ⑤ Disposal — sub-task timeout (alongside /stop · cron pause)
curator:
  enabled: true             #    auto-archive unused skills

Delegation is not handed over all at once either. At first, entrust only the narrowest, most easily reversible tasks (read-only work like the integrity check), watch the results for a few days, then widen to the next seat. Choosing the nightly integrity check as the first task in §23.2.4 was for the same reason — it only reads, so the worst case is a single wrong report, and the originals are untouched.


23.2.7 Adoption Progress and Incremental Stages (as of June 2026)

As of this chapter's revision, adoption has entered the settling-in phase. The native Windows build (v0.16.0) is installed, and model API key registration via hermes setup is done. I brought the first seat online and checked the five safeguards one by one against their actual config keys, and I am now running real autonomous tasks to get them into muscle memory. To be honest about it: I am validating on my personal PC first, not the company PC — company adoption is deferred until the safeguards are thoroughly worn in at home. This is less caution than the PC separation principle. You do not unleash an unvalidated autonomous tool on team data.

Period Activity Gate
Month 1 Install Hermes (native Windows v0.16.0) + hermes setup + first task Are all five safeguards on?
Months 2–3 Expand to two or three seats (meeting-note classification, build capture analysis) Log check at each widening of delegation
Months 3–6 Company review — decide based on personal-PC validation results Zero unattended-operation incidents confirmed
Months 6–12 Team-level adoption Safeguards settled in as team rules

The temptation to skip stages is the most dangerous part. Jump straight from month 1 to month 6 (team adoption), and the safeguards get released while they are still one person's habit, not yet learned as team rules. The answer is to stop once at the end of each stage and check the five safeguards. Going in a way that stays reversible matters more than going fast.


23.2.8 Five Common Misconceptions

"The agent replaces the human" is the most common misconception. The worked transcript in §23.2.4 shows the opposite — the agent stopped at a single column mapping, and a human filled in that judgment. The core decisions of game design still belong to the human; what the agent takes is the tedious part of repetition and analysis.

The expectation that "one install and everything is automatic" is dangerous too. The first month or two actually demand more of your hands, not less. Column-name mappings, permission scopes, and cost limits have to be tuned per task, and until that tuning settles, a human reviews every output.

The verdict that "Claude Code is now obsolete" is wrong. The two occupy different hours. Daytime precision decisions go to Claude Code; nighttime unattended repetition goes to the agent. The /check Cascade of §23.1 did not disappear — a night lane was added beside it.

The notion that "it's open source, so it's free" is half right. The body may be free, but model API call costs accrue all the same. That is why the cost limits in config.yamlagent.max_turns, compression, and the like — are both a safeguard and a household ledger.

Finally, the expectation that "the agent will handle even the complex, risky work" is the most dangerous of all. The higher the risk of a task, the more firmly it stays under human control. What you hand to the agent starts with the simple and the easily reversible. Delegation widens only as far as trust has accumulated.


23.2.9 Into the Next Chapter

If the Wrapper, Cascade, and Junction of §23.1 are the summit of Claude Code operations, this chapter's Hermes lays one more lane — a night lane — on top of that operation. The picture of a daytime tool and a nighttime tool sharing the same desk: that is both the present as of 2026 and the skeleton of the near future.

The next chapter is tool curation for game designers. What goes into the 12 slots, what gets culled by skill_audit_score — the curation criteria this chapter only brushed past get unpacked into concrete tool recommendations.


Key Takeaways

Next Chapter Preview


Try It Yourself

setup 1. Download and install the Hermes native Windows installer (if you prefer Linux, the path of wsl --install followed by installing inside it is still there). 2. Create a working folder on a fast local disk and copy the data sheets you want to check into it (do not wire a network drive or an SVN working folder directly in as the workspace). 3. hermes setup → enter your model API key → confirm the workspace path and initial permission values. 4. Turn on the five safeguards in %LOCALAPPDATA%\hermes\config.yaml: permission approval (approvals.mode: manual · command_allowlist · cron_mode: deny), cost limits (agent.max_turns · terminal.timeout · tool_loop_guardrails), checkpoints (checkpoints.enabled), the log path (logs.path), and learn the stop procedures (/stop · /rollback).

prompt - Throw the goal one level more abstract: not "do this for me" but "have this result ready." - Spell out the target, the work, and the save location as numbered items, and always end with one line: "If you get stuck, or a column/format differs from what's expected, stop and report. No guessing." - Pick something easily reversible for the first task, like a read-only integrity check.

verify - Do not trust the first output as is; check the spot where the agent stopped (column mapping, format mismatch) yourself. - If the halt was right, re-request with the mapping made explicit; if it was wrong, restate the constraints. - Cross-check one or two failure entries in the generated report against the original sheet to confirm the agent's judgment, and only then move it to the nightly schedule (cron 23:00).

Solo Scale-Down

If you want to get a feel for the agent before installing Hermes, you can run a scaled-down version inside Claude Code using background execution.

Part 23 · Chapter 3. Tool Curation — Cutting Unused Tools with Data

During a quarterly retrospective, I opened my global skills folder. Counting line by line, there were 19 wrappers. I had clearly decided to run with 12, and had run it that way for a year — yet somewhere along the way seven more had attached themselves. What made it more absurd: for half of them, I couldn't tell from the name alone what the tool even did. migrate-legacy-enum. What was this again? When did I last use it?

I couldn't remember. And as long as I rely on memory, that question can never be answered. So I decided to look at logs instead of memory. Tool curation should not be cutting by taste — it should be cutting by a number: "how many times did I call this tool last quarter?"

This chapter is a record of how I pull that number out automatically, how I use it to cut tools, and how I keep tools from proliferating in the first place.


23.3.1 Tools Multiplying Is a Natural Phenomenon

Before talking about curation, one thing has to be acknowledged: unless you stop them, tools will multiply. It is not weak willpower. On every task, making one small script "just to handle this one quickly" is the rational choice. Stack that rational choice up a few dozen times and you get an irrational pile.

The structure I run on Project A is 12 global wrappers pointing, via junctions, to the 48 actual tool bodies in the workspace. The global side stays light; the heavy bodies live in the SVN-managed workspace. The structure itself was covered in §23.1. The problem is that this number 12 refuses to sit still.

Look at what grows alongside each new tool, and it becomes clear why this has to be stopped.

Adding 1 tool → 4 costs that grow with it New tool +1 Context token usage ↑ Choice fatigue ↑ Maintenance surface ↑ Feature overlap risk ↑ The tool is +1, but the cost is +4. This is why curation is a subtraction job.

The first cost in particular — context token usage — has become a sharper cost in the era of AI tools. As global wrappers multiply, so do the tokens the AI spends every session reading "the list of tools I can use." Reading the descriptions of 19 tools eats into the context that should go to the actual work. That is why Project A's sync_skills.py has a --cleanup option that automatically clears out wrappers whose junctions are broken or whose bodies have disappeared. It is closer to hygiene work for protecting the token budget.

But --cleanup only catches "broken" tools. A tool that is perfectly alive but that nobody uses slips through. To catch those, you need usage-frequency data.


23.3.2 skill_audit_score — Measuring Usage Frequency from SVN Logs

The core idea is simple. Every skill and tool in the workspace lives in SVN. And every time a tool is used, the artifacts it produces (sheets, documents, relation-map HTML, and so on) get committed to SVN. In other words, the SVN log carries the traces of which tools actually did work.

So I built a small measurement script called skill_audit_score. As the name says, it assigns each skill an "audit score." When I built it, I didn't write all the code from scratch — I explained the measurement logic to the AI and got a draft. I'm leaving that worked transcript as is: not summarized, and including the part where it stepped wrong.

The Prompt (in Full)

My global skill wrappers are fixed at 12, but they have swollen to 19 at some point. For the quarterly retrospective I want to pick out the "unused tools" with data. Build a measurement script on the following assumptions.

Claude's Raw Output (Verbatim)

Sounds good. One thing to flag before we proceed, though. There is no guarantee that commit messages always contain the tool name. When people write commit messages freely, a message like "edited sheets" tells us nothing about which tool was involved. So I will run two kinds of matching in parallel.

  1. Keyword matching on tool names and aliases in the commit message (loose)
  2. Matching whether a changed file path falls inside that tool's folder or matches the tool's artifact pattern (e.g., *_relation.html) (strict)

Since #2 is more trustworthy, I'll give it a higher weight. Here is the draft.

# skill_audit_score.py  (Claude draft — pre-verification)
import subprocess, sys, datetime, xml.etree.ElementTree as ET
from collections import defaultdict

def svn_log_xml(repo_path, limit=2000):
    out = subprocess.run(
        ["svn", "log", "--xml", "-v", "-l", str(limit), repo_path],
        capture_output=True, text=True, check=True)
    return ET.fromstring(out.stdout)

def score_skills(repo_path, skills, window_days=90):
    root = svn_log_xml(repo_path)
    now = datetime.datetime.now(datetime.timezone.utc)
    hits = defaultdict(list)   # skill -> [days_ago, ...]
    for entry in root.findall("logentry"):
        date = datetime.datetime.fromisoformat(
            entry.find("date").text.replace("Z", "+00:00"))
        days_ago = (now - date).days
        if days_ago > window_days:
            continue
        msg = (entry.find("msg").text or "").lower()
        paths = [p.text for p in entry.findall("paths/path")]
        for sk in skills:
            name = sk["name"].lower()
            in_msg = name in msg or any(a in msg for a in sk.get("alias", []))
            in_path = any(name in (pp or "").lower() for pp in paths) \
                      or any(pat in (pp or "") for pp in paths
                             for pat in sk.get("artifact", []))
            if in_path:
                hits[sk["name"]].append((days_ago, 2.0))   # strict match, weighted
            elif in_msg:
                hits[sk["name"]].append((days_ago, 1.0))   # loose match

    rows = []
    for sk in skills:
        h = hits[sk["name"]]
        if not h:
            rows.append((sk["name"], 0, None, len(h)))
            continue
        # recency weight: 1.0 when recent, down to 0.1 at 90 days out
        recency = sum(w * (1 - 0.9 * (d / window_days)) for d, w in h)
        last_used = min(d for d, _ in h)
        # last-used penalty
        stale_penalty = max(0, (last_used - 14)) * 0.5
        score = max(0, min(100, recency * 8 - stale_penalty))
        rows.append((sk["name"], round(score, 1), last_used, len(h)))

    rows.sort(key=lambda r: r[1])   # ascending score = candidates first
    return rows

if __name__ == "__main__":
    REPO = sys.argv[1]
    SKILLS = [
        {"name": "relation-map-gen", "alias": ["relation", "관계도"],
         "artifact": ["_relation.html"]},
        {"name": "schema-doc", "alias": ["schema", "명세서"],
         "artifact": ["_schema.md"]},
        {"name": "gdd-gen", "alias": ["gdd"], "artifact": ["_gdd.md"]},
        {"name": "table-creator", "alias": ["table", "테이블"], "artifact": []},
        {"name": "excel-reader", "alias": ["xlsm", "vba"], "artifact": []},
        {"name": "migrate-legacy-enum", "alias": ["enum", "migrate"],
         "artifact": []},
    ]
    print(f"{'skill':22}{'score':>7}{'last_used(d)':>14}{'hits':>6}")
    for name, score, last, hits_n in score_skills(REPO, SKILLS):
        print(f"{name:22}{score:>7}{str(last):>14}{hits_n:>6}")

Human Verification — What I Ran and Rejected

I ran the script against my actual SVN working copy. The first output looked wrong. table-creator and excel-reader both had empty artifact patterns, so they were only caught when the words "table" / "xlsm" happened to appear in a commit message. But these two tools produce plain .xlsm files as their artifacts, which no pattern can distinguish. As a result, both came out with abnormally low scores — even though they are tools I use every day.

This is where I made an important call. A low score must not mean an automatic cut. A human has to separate whether the score is low because the tool genuinely goes unused, or because the measurement fails to catch it. The numbers the AI produces only narrow the candidates; the final decision belongs to a human.

So I went back to the AI.

The Follow-Up Prompt

Tools with an empty artifact pattern have scores we can't trust, so add a confidence column to the output. Mark tools that never had a single artifact match as confidence=LOW and exclude them from the automatic curation candidates. Group the LOW tools separately under "unmeasurable — manual review."

This follow-up split the output into two buckets: tools that can be cut on a trustworthy score, and tools whose measurement is too weak and that a human has to inspect directly. The shape of the actual run came out roughly like this (scores are real measurements from my working copy; some tool names are anonymized).

skill audit_score last_used (days ago) confidence Verdict
relation-map-gen 71.4 2 HIGH Keep
schema-doc 58.9 5 HIGH Keep
gdd-gen 22.1 31 HIGH Watch
migrate-legacy-enum 0.0 not measured HIGH Curation candidate
table-creator 4.2 1 LOW Manual review → keep
excel-reader 6.0 1 LOW Manual review → keep

migrate-legacy-enum scored 0 with HIGH confidence. That means in 90 days, neither the tool's folder nor its artifacts appeared in a single commit. Digging through my memory: it was a job that should have stayed one-off — migrating a legacy enum once last year, done — that I had codified into a skill. This is exactly the tool to cut. Conversely, table-creator and excel-reader scored low but with LOW confidence, and their last use was one day earlier. The measurement simply couldn't see them; in reality they are used daily. They must not be cut.

Note: the scoring formula in the table above (recency weight × 8, stale penalty) is a set of values I tuned to my own working copy. With different SVN commit habits and artifact patterns, the coefficients change too. The essence of this tool is not the absolute score but the "relative ranking among tools" and the "confidence split."


23.3.3 The Curation Cycle — From Measurement to Retirement

skill_audit_score is only a measurement tool. Tools actually get cleaned up only when the measurements are slotted into the quarterly retrospective and run through a full cycle. That cycle is the following.

flowchart TD
    A[Quarterly retrospective starts] --> B[Run skill_audit_score
parse 90 days of SVN log] B --> C{Confidence verdict} C -->|HIGH| D{Evaluate audit_score} C -->|LOW| E[Move to manual-review queue
check last-used date directly] D -->|High score| F[Keep] D -->|Middling, trending down| G[Watch — remeasure next quarter] D -->|Zero or rock bottom| H[Confirm curation candidate] E --> F E --> H H --> I{Replaceable?} I -->|Absorb into wrapper| J[MECE augmentation of existing tool
§23.1 wrapper policy] I -->|Retire completely| K[sync_skills.py --cleanup
remove junction + archive in SVN] J --> L[Confirm 12 slots recovered] K --> L L --> A classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; class B,C,K code; class A,D,E,I human; class F,L pass;

What matters is distinguishing the cycle's two exits. A tool that scores 0 does not get deleted unconditionally. If the work itself has disappeared, it goes to full retirement (--cleanup); if the work is still needed but not often enough to justify a separate tool, it gets absorbed into an existing tool. The latter is the MECE augmentation of §23.3.4.

Even on retirement, the code remains in SVN history. Only the junction and the global exposure are withdrawn — the code itself is not erased forever. If that job comes back six months later, you restore it from SVN. It is this "reversible" safety net that lets a human cut boldly.


23.3.4 Curbing MECE Sprawl — Ask Before You Build

Better than measuring and cutting is not building in the first place. If skill_audit_score is after-the-fact cleanup, the MECE wrapper policy is up-front restraint.

MECE stands for Mutually Exclusive, Collectively Exhaustive — no overlaps, no gaps. Every time I want to build a new tool, I throw those four letters at it. Does the new tool overlap an existing tool (an ME violation)? Or does it genuinely fill an empty area (a CE contribution)? Project A's wrapper policy forks two ways from here.

Situation Policy Result
The new job overlaps an existing tool's area Augment the existing tool first Add the feature to the existing wrapper's body; no new slot used
The new job is clearly a different area Allow a new wrapper Assign one of the 12 slots to the new tool (paired with a candidate to remove)

The key is that the default is augmentation. Building a new tool is the exception. To justify that exception, you have to prove that "no existing tool can do this job." This one default is the real reason the tools that had swollen to 19 came back down to 12.

This also connects to the cascade in §23.1. A cascade like check is the result of bundling what used to be four separate checking tools into a single call. Instead of keeping four separate wrappers, the MECE view said "this is all one area called checking" and absorbed them into one. The tool count went down; the functionality stayed the same. That is the model case for augmentation.

The AI assistant is both the hazard and the remedy here. It is a hazard because asking the AI to "make me a script that handles this" produces a new tool far too easily. In an environment where one click births one tool, without MECE discipline a tool graveyard forms in no time. It is a remedy because if you hand the AI the policy first, the AI will itself suggest "this would be better as an option added to the existing relation-map-gen." The AI that builds your tools must be handed the curation discipline along with them.


23.3.5 How Not to Be Fooled by the Score — Limits of Measurement

What I learned most while running this chapter's tool is that you must not put blind faith in the measurements. skill_audit_score looks at exactly one signal: the SVN log. So there are things it structurally misses.

In short, this tool is not a "deciding tool" but a "candidate-narrowing tool." It looks across all 19 at a glance and tells you, in one second, which ones to suspect. Verifying that suspicion and making the cut stays the human's job. Measurement does not replace the human — it only points at where the human should look.


Try It Yourself — One Cycle of skill_audit_score

This is the procedure for turning the tool curation cycle through one full lap yourself.

setup 1. Confirm that your workspace's skills and tools are under version control (SVN/Git). Their artifacts must be getting committed to the same repository. 2. Make a list of the tools to measure. For each tool, record name, alias (aliases that appear in commit messages), and artifact (artifact file patterns, if any). Leave artifact empty for read-only tools that have none.

prompt (to the AI)

Build a tool-usage-frequency measurement script on the following assumptions. (1) Each tool leaves traces in the [version control system] log as artifact commits. (2) Parse the last 90 days of the log and count the commits each tool was involved in. (3) Produce a 0–100 score from recency weighting plus a last-used-date penalty. (4) Mark tools with zero artifact-pattern (artifact) matches as confidence=LOW, exclude them from the automatic candidates, and split them out for manual review. (5) Output is a table in ascending score order — low scores are the curation candidates. Use only the standard library; take the repository path as an argument.

verify 1. If a tool you use every day lands at the top of the table (low score), the measurement is wrong. Check that tool's confidence — LOW is normal (unmeasurable); if it is HIGH yet scoring low, inspect the alias and artifact settings. 2. Confirm only tools with score 0 + confidence HIGH as curation candidates. Cross-check the last-used date against your memory, and let a human judge whether the tool is truly dead. 3. Send each candidate down one of two paths: "full retirement" or "absorption into an existing tool." Retirement withdraws only the junction; the code stays in the repository. 4. Finally, count whether your 12 slots (or whatever cap you set) have been recovered.

Solo Scale-Down

If you are a solo developer with only 6–8 tools and no SVN, scale it down like this. Git is plenty for version control. Pull the changed file paths with git log --since="90 days ago" --name-only, grep once for your tool folder names, and out comes "which tools worked recently." You don't even need to build the scoring script. The point is not numeric precision — it is the single habit of looking at logs instead of memory. Once a quarter, pull "the tools untouched in the last 90 days" from the git log and stare them down. Those five minutes prevent the tool graveyard.


Key Takeaways

Next Chapter Preview

Part 23 · Chapter 4. The Puzzle Game I Built Alone — A Critter Sort Field Report

One Saturday afternoon, my wife was playing a color-matching puzzle on her phone. It was a game called Yarn Fever, where you sort tangled skeins of yarn into baskets of the same color. Every time a round ended she would say, "Same thing again," and close it. With nothing new coming, she got bored fast.

The thought that hit me at that moment was simple. That loop has proven addictiveness, and the mechanic itself is not subject to copyright. Swap in an animal theme, stamp out levels procedurally without end, and the "same thing again" problem disappears. Build it alone, as HTML 3D that runs right in the browser, and my wife doesn't even need to install anything on her phone.

The problem is that I am not a graphics engineer. Twenty-four years as a game designer, but I had never written a shader with Three.js. So this chapter is the actual record of getting one game running in a few days, alone, with AI. It is also a record of separation: I used the same tools as my company MMORPG (Project A hereafter) while mixing in not a single line of its domain content.

The actual game lives in the critter-sort/ repository, and git tags v0.1–v0.3 preserve three days of decisions. This is not a polished-up case study; I quote that repository as is.


23.4.1 Reverse Engineering by Prompt — And Losing the Signature

The first thing I did was break the original down into words and throw it at the AI. The first prompt went like this.

Prompt (v0.1 kickoff): "I want to adapt the core loop of a casual puzzle called Yarn Fever to an animal theme, built with Three.js + Vite. The loop goes like this: you sort tangled color-clumps into bins of the same color, and if you overflow the temporary slots, it's game over. Let's make animals the things being sorted — tap a tangled pile of animals and they go to a nest of the same color. Write the logic as a pure JS state machine independent of Three.js so it can be tested headless. Also add procedural infinite levels (seed-based)."

The AI followed faithfully. It split the folder structure into game/ (pure logic) and render/ (Three.js), wrote state.js, rules.js, and generator.js first, then put the board on screen with colored-box placeholders. v0.1 was running not in days but in a single session.

But the moment I played it myself to show my wife, something felt off. It had become a generic match puzzle where animals hop across the screen into baskets. The feel of the original was gone. Yarn Fever's identity was not "sorting" — it was the tactile sensation of untangling skeins by hand and the act of rotating the screen to check colors hidden from view. In my prompt I had flattened that into a generic sort — "tap an animal and it goes to a nest" — and the AI had simply been faithful to the definition I gave it.

This is the first trap of reverse engineering. Summarize the original and the signature evaporates, because a summary preserves the surface, not the essence.

Let me make one thing clear here. The AI did not give a wrong answer. My prompt was wrong. So the next cycle was not a code fix but a correction of the definition.


23.4.2 The Spine — The Correction Cycles That Revived Rotation and Unraveling

Right after seeing v0.1, I respecified the core in two passes. The first correction targeted the action; the second targeted the feel.

Prompt (correction 1 — the action): "This has turned into a plain match puzzle. Yarn Fever's identity is 'rotating the 3D view to check hidden colors and deciding the order to unravel.' Don't scatter the board flat — pile it into a tall mound in the center so you have to rotate to see the colors hidden behind. Add rotation with OrbitControls. But rotation and tap are going to conflict — how will you separate them?"

Claude's output (raw, no summarizing): "You're right. The rotation gesture and the tap are the same pointer events, so they conflict. I propose splitting them by distance and time thresholds. Record the coordinates and timestamp on pointerdown; on pointerup, if the travel distance exceeds 9px or the press duration exceeds 400ms, treat it as a rotation gesture and ignore the tap. Only below that do we pick an animal with the raycaster and call state.tap(). I'll clamp OrbitControls' azimuth to about ±54° so the pile doesn't drift off screen. One caveat: 9px may be too sensitive on mobile, so it needs checking on a real device."

That 9px/400ms threshold still lives in src/main.js today. No-frills code you can verify directly.

// src/main.js — drag = rotate (OrbitControls) / short tap = unravel yarn
canvas.addEventListener('pointerup', (ev) => {
  if (!down || state.status !== 'playing') { down = null; return }
  const dist = Math.hypot(ev.clientX - down.x, ev.clientY - down.y)
  const dt = performance.now() - down.t
  down = null
  if (dist > 9 || dt > 400) return // treated as a rotation gesture
  const rect = canvas.getBoundingClientRect()
  const x = ((ev.clientX - rect.left) / rect.width) * 2 - 1
  const y = -((ev.clientY - rect.top) / rect.height) * 2 + 1
  const id = view.pick(x, y, camera)
  if (id !== null) state.tap(id)
})

The action was back. But on a second playthrough it still didn't feel like "unraveling a skein." Animals simply moved to nests; there was no tactile sense of something coming undone. Second correction.

Prompt (correction 2 — the feel): "The rotation is better. But there's still no feeling of 'unraveling yarn.' On tap, I want the tangled strands wound around the animal's body to come undone and disappear, and that yarn to wind its way into the basket. For the animal look, not sleek GLB birds — chubby knitted plushies (amigurumi). Cozy details like sleepy eyes and blushed cheeks. If external assets don't fit, drop them and generate procedurally."

Claude's output (raw): "I'll implement this as two separate pieces. (1) Tangled yarn: wrap two strands of slightly darkened same-color TorusGeometry around each animal's body, and on tap shrink their scale to 0.01 so they unravel and vanish. (2) Yarn wind-up: draw the yarn along the animal-to-basket path with TubeGeometry and tween it winding into the basket. The plushie look goes procedural — I'll procedurally generate a knit bump texture on MeshStandardMaterial for the knit feel, flat boxes for the sleepy eyes, translucent pink spheres for the cheek blush. The existing GLB starters (three birds) clash with the knit tone, so I'll remove them but keep the animal-pack load pipeline alive as a fallback. One thing to flag: procedural animals can't match GLB for species variety — it's safer to validate with one species first and then add more."

That final warning ("can't match GLB for species variety" — GLB is glTF Binary, the off-the-shelf 3D model file format you download and use) was the seed that led straight to v0.3. The AI named the next limitation first, and I took it as the next milestone.

Verification was two-stage every time: headless, to confirm the logic hadn't broken (0 errors), then hands-on in the browser, rotating and tapping to check the feel. The v0.2 commit message pins that verification down: "headless verification: rotation, yarn unraveling, auto-clear all normal, 0 errors."

The two tangled strands remain in src/render/pieces.js like this.

// src/render/pieces.js — two loose yarn strands wrapped around the body (same color, slightly darkened)
const strandMat = new THREE.MeshStandardMaterial({ color: darken(hex, 0.7), roughness: 1 })
const strands = []
const orient = [[0.5, 0.2, 0.0], [1.25, 0.0, 0.6]]
for (let i = 0; i < 2; i++) {
  const s = addMesh(g, G.torus, strandMat, [0, byo + 0.02, 0], Math.max(bx, bz) + 0.02, orient[i])
  strands.push(s)
}
g.userData.strands = strands  // on tap, view.js unravels these strands and fades them out

The lessons from this, one line each:


23.4.3 Three Days of Decision History — Reading the Corrections Through Git Tags

In words it's just "I fixed it twice," but the git history recorded when and in what form each correction landed, with exact timestamps. In solo development this stands in for a retrospective. Even with no teammates, the commits testify to "why it ended up this way."

Commit Time (2026-05-30) What Changed Signature Status
2b2e3bc v0.1 14:43 Yarn Fever reverse engineering, pure logic + placeholders, 60/60 solver pass Missing (flattened into a generic sort)
70a0117 v0.2 15:11 Rotation (OrbitControls ±54°) + tap/drag separation + yarn unraveling + amigurumi Restored (core redefined)
160663c snapshot 15:31 5 v0.2 gallery snapshots + README gallery
59b0baf v0.3 15:55 8 procedural amigurumi species + vivid candy palette Reinforced (species variety secured)
c5b9a1b handoff 16:20 NEXT_SESSION session handoff pointer

The body of the v0.2 commit message pinned down the decision itself: "Corrected the game's identity from 'animals hopping' to 'rotate the view, unravel cute skeins (knitted plushies), and sort them into same-color baskets.'" It is the record of a game's identity dying once and coming back to life inside an hour and a half.

One detail worth noting. Look at v0.2's git show --stat and the three starter GLB birds (Flamingo, Parrot, Stork) were deleted wholesale — "because the art clashed with the knit tone." Free external assets don't get used just because they're free; if the tone doesn't fit, they go. That is an aesthetic gate a human applied, not the AI.

public/assets/animals/pack_starter/Flamingo.glb  | Bin 77428 -> 0 bytes
public/assets/animals/pack_starter/Parrot.glb    | Bin 97024 -> 0 bytes
public/assets/animals/pack_starter/Stork.glb     | Bin 76852 -> 0 bytes

23.4.4 Procedural Generation in Practice — Eight Amigurumi Species and Infinite Levels

The homework v0.2 left behind was "procedural animals can't match GLB for species variety." v0.3 solved it. Without adding a single external asset, I stamped out eight animal species in code.

The core is the SPECIES table in src/render/pieces.js. Each species defines its body proportions, head, ear type, snout, and eye shape as parameters, and one function reads those parameters and assembles the mesh.

// src/render/pieces.js — per-species silhouette parameters
const SPECIES = {
  cat:      { body: [0.5,0.46,0.48,0.04], ears: 'cat',   snout: 0.13, tail: 'cat',  eyes: 'sleepy' },
  bear:     { body: [0.52,0.5,0.5,0.03],  ears: 'bear',  snout: 0.16, tail: 'none', eyes: 'round' },
  bunny:    { body: [0.46,0.5,0.46,0.02], ears: 'bunny', snout: 0.12, tail: 'puff', eyes: 'round' },
  fox:      { body: [0.5,0.44,0.48,0.04], ears: 'fox',   snout: 0.2,  tail: 'fox',  eyes: 'sleepy' },
  capybara: { body: [0.58,0.5,0.56,0.02], ears: 'tiny',  snout: 0.22, tail: 'none', eyes: 'sleepy' },
  pig:      { body: [0.54,0.5,0.52,0.03], ears: 'pig',   snout: 0.1,  nose: true,   eyes: 'round' },
  frog:     { body: [0.56,0.4,0.54,0.05], ears: 'none',  snout: 0.1,  topEyes: true, eyes: 'none' },
  chick:    { body: [0.42,0.44,0.42,0.05], ears: 'none', beak: true,  tail: 'none', eyes: 'round' },
}
export const SPECIES_IDS = Object.keys(SPECIES)  // 8 species

Ear shape alone splits the silhouettes. Cats and foxes get pointed cones, bears get round spheres, bunnies get elongated spheres, pigs get cones folded forward. Frogs get eyes that pop out above the head (topEyes); chicks get a beak (beak). These small branches create the distinguishability of all eight species. Zero external assets, one file of code.

But procedural generation has a trap. Whether plausible-looking code actually produces eight identifiable species is something you cannot tell from the code alone. So verification was two-stage again: headless, to confirm all eight species generate without errors, then the web-screenshot skill (headless Chrome) to capture actual renders and confirm by eye that the eight are distinguishable. The result is in DEVLOG v0.3: "silhouettes distinguished by ears/snout/nose/beak/tail/eyes. Zero external assets, knit tone fully unified."

One Seed Determines the Entire Board

The infinity of levels is the seeded RNG's responsibility. generator.js turns the level number into a seed with a Knuth multiplicative hash, then draws deterministic random numbers with mulberry32. The same level number always yields the same board.

// src/game/generator.js
export function generateLevel(level, animalPool = null) {
  const seed = (level * 2654435761) >>> 0  // Knuth multiplicative hash
  const rng = makeRng(seed)
  const { C, K, groupsPerColor, M, T } = levelParams(level)
  const colors = rng.shuffle(COLORS).slice(0, C)
  // ...
  for (const color of colors) {
    const count = K * groupsPerColor  // always a multiple of K → divides exactly into nests (solvability guaranteed)
    // ...
  }
}

One line here guarantees the game's fairness. Because the animal count per color is forced to be always a multiple of K (the number of animals that completes a nest, 3), every board divides exactly into nests. An unsolvable level cannot occur in the first place.

Do All 60 Levels Actually Clear? — greedySolve

"Solvable by design" is not a proof. I put a greedy solver for verification into rules.js, and test-logic.mjs auto-plays all 60 levels to confirm, every time, that they all actually clear. Here is the measured output from rerunning it just now, while writing this chapter.

$ node scripts/test-logic.mjs
[solver] 60/60 levels cleared

[difficulty curve] (C=colors, K=to complete, groups, M=nests, T=tray, total animals)
  Lv 1: C=3 K=3 grp=2 M=3 T=7 total=18
  Lv 8: C=4 K=3 grp=3 M=4 T=6 total=36
  Lv12: C=5 K=3 grp=3 M=4 T=5 total=45
  Lv20: C=5 K=3 grp=3 M=4 T=4 total=45

[button-mashing play] loss rate on random taps (checking that difficulty exists)
  Lv 1: random loss rate 0%
  Lv12: random loss rate 1%
  Lv20: random loss rate 3%

This test proves two things at once. The greedy solver clearing 60/60 means every level is solvable (the difficulty is not impossible), and the random-tap loss rate climbing from 0% to 3% as levels rise means the difficulty is real (if mashing at random clears everything, it isn't a game). The difficulty curve of the tray narrowing from 7 slots to 4 is measured as a loss rate.

Let me be honest here. The 3% random loss rate is the loss rate of a button-mashing bot, not a human's perceived difficulty. A human checks colors in advance by rotating, so their loss rate is lower. This number is a directional proof that "the difficulty is not zero," not a claim that my wife loses 3% of the time. Human-perceived difficulty was still unmeasured as of v0.3, and I left it in NEXT_SESSION as "collect wife's play feedback (top priority)."

The Procedural Generation Pipeline

flowchart TD
  L["Level number N"] --> H["Knuth hash
N × 2654435761"] H --> S["mulberry32(seed)
deterministic RNG"] S --> P["levelParams(N)
derives C·K·M·T"] P --> G["generateLevel
per color = multiple of K"] G --> B["Board (tangled pile)"] B --> R["createCritter
8 amigurumi species + knit shader"] G --> V["greedySolve
60/60 verification"] V -->|"0 errors"| OK["Clear guaranteed"] R --> SC["web-screenshot
visual check: 8 species distinguishable"] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; class H,S,P,G,R,V,SC code; class L,B data; class OK pass;

This flow — starting from a seed and branching into parameters, board, meshes, and verification — is the structural answer to the original problem: "nothing new, so it gets boring."


23.4.5 What If a GLB Comes In? — Auto Scale and Fallback

The eight procedural animals are the fallback for when there is no GLB. I kept the animal-pack pipeline alive so that if I get real amigurumi GLBs later, those take priority. Drop the GLBs into the folder, run npm run scan, and that's it.

The problem is that every GLB comes in a different size. One model is 0.5 units, another 200. Match the scale by hand and adding an animal pack becomes labor. So scan-packs.mjs reads the GLB's bounding box and automatically computes the scale for the target height (0.95 units).

// scripts/scan-packs.mjs — auto-derive scale/yOffset from the GLB bounding box
const maxDim = Math.max(max[0]-min[0], max[1]-min[1], max[2]-min[2])
const scale = +(TARGET_H / maxDim).toPrecision(3)        // TARGET_H = 0.95
const yOffset = +(-((min[1] + max[1]) / 2) * scale).toPrecision(3)

And assets.js quietly falls back to the procedural animals if packs.json is missing or fails to load.

// src/render/assets.js
createAnimal(species, hex) {
  const entry = this.models.get(species)
  if (!entry) return createCritter(hex, species)  // procedural amigurumi fallback
  // ... GLB clone + color tinting
}

These two lines guarantee "GLB if available, code animals if not" with zero downtime. I can drop a new GLB pack in while my wife is playing and the game never stops.


23.4.6 Alone but Working Like a Team — How I Used AI

On this project I was one game designer, but the work ran across multiple roles. AI filled those roles. The point was not "it writes the code for me" but it fills in where I am weak.

Where I Am Weak What the AI Did The Gate the Human (Me) Kept
Three.js shaders Procedural knit bump texture, TubeGeometry yarn effects Does the tone fit? (the decision to delete the three GLB birds)
Resolving input conflicts Proposed the 9px/400ms thresholds Confirming the feel on a real mobile device
Regression safety Automatic 60/60 verification with greedySolve "Difficulty is real" is defined by a human
Predicting the next limitation Warned that "procedural animals are weak on species variety" Adopting it as the v0.3 milestone

Visual verification in particular was the weak link of solo development. Code that runs and "eight species distinguishable by eye" are different problems. So I borrowed the web-screenshot skill (spin up the dev server in headless Chrome, capture screenshots, and report console errors) straight from my company workflow. Even without the claude-in-chrome extension, I could check the mobile-viewport render (iPhone 15 Pro portrait, 393×852) with my own eyes.

Here the most important principle is at work: borrow tools from the company, borrow zero domain content.

This separation is verified with grep. My memory log records "company project domain content borrowed: 0 instances (verification grep PASS)." Critter Sort's colors are pink, mint, and yellow; its animals are cats, bears, and bunnies. The domain vocabulary of Project A (the company MMORPG) appears nowhere in this repository.

Why separate this thoroughly? To block two accidents at once: the legal accident of company IP leaking into a personal hobby, and the context contamination of MMORPG domain atoms being wrongly injected into puzzle work as noise. Tools flow, content is blocked — that line is what healthy separation looks like.


23.4.7 Wrap-Up — The System Works Even Solo

Critter Sort is a small game. Three days, 5 commits, 8 animal species, 60 levels. And yet the methods I used at the company worked unchanged at 1/1000th the scale.

The biggest lesson was the failure in the first section. In v0.1 I killed the game's identity once, then revived it with two corrections. In solo development with no teammates, what testified to that death and resurrection was the git commits. Without that retrospective record, a month later I would have forgotten "why did v0.2 rip everything out?"

Part 24, next, covers how to harden this kind of decision history into governance for large teams and long-term live ops.

This chapter has been the record of a journey that started with a puzzle my wife closed out of boredom and ended with a game built alone landing back in her hands. It confirmed that a system is a matter of discipline, not scale.


Try It Yourself — One Step You Can Take Today

This step is about getting the core loop of a casual game you love running with AI — without losing the signature.

setup — On a machine with Node installed, create an empty folder: mkdir my-puzzle && cd my-puzzle.

prompt — Throw this at the AI. The key is to "spell out the signature instead of summarizing."

"I want to adapt the core loop of [game name] to [theme]. This game's signature is [write the feel in one line — e.g., 'the tactile loop of rotating the view to reveal what's hidden and untangling it']. Do not flatten this into a generic match puzzle. Write the logic separate from the render so it can be tested headless."

verify — Play the first result yourself. Ask: "Is the signature I wrote down still alive?" If it isn't, rewrite the definition — not the code — and ask again. That is what I did going from v0.1 to v0.2.

Scaled-Down Version for Solo and Hobbyist Readers

You don't need an engine or procedural generation. Write "this game's signature in one line" on a piece of paper, have the AI build a prototype, then play it yourself and check just one thing: did that line survive? If it died, rewrite the line more concretely. The habit of guarding a one-line signature — that alone is enough to avoid the first trap of reverse engineering.

Part 24 · Ops Deep

24.1 The Verification System — Catching Consistency, Links, and Staleness in Code

Right after Monday morning standup, team member A from the data team sent me a screenshot over the team messenger. It was a QA report: the description of a material item was showing up empty in the in-game shop. Thirty minutes of tracing later, the culprit surfaced. Two weeks earlier, someone had renamed the item to 재료_목재_상 ("material_wood_high") in a design document, but the reference in the data sheet still pointed to the old name, 재료_목재_A ("material_wood_A"). The document was updated, the sheet was not, and the link between them had quietly snapped. Nobody lied, yet the game was printing a lie.

Incidents like this multiply exponentially as documents pile up. Human eyes cannot track the cross-references of 50 documents at once. So I delegate verification to code. This chapter covers a system where scripts, not people, check the consistency of documents, data, and links. The core is three things — source consistency (the _source_map.tsv audit), link integrity (wikilinks), and stale detection (catching references that have gone rotten with age).


24.1.1 Why Broken Links Are Silent

Documents and data live by pointing at each other. A design doc references an enum, the enum references a data sheet, and the sheet references a decision in yet another design doc. Manage this web by hand, and every time one node changes, a human has to remember every reference pointing at that node and chase them all down. Memory fails.

Broken links are dangerous because they throw no errors. In code, the compiler stops you when you reference a variable that doesn't exist. But a wikilink written as [[재료_목재_A]] in a document just stays there as ordinary text after its target disappears. It doesn't turn red. The game builds, ships, and only after a player sees the empty description does anyone notice.

So the verification system's first job is making visible what human eyes cannot see. Pull consistency violations out into text output, tie that output to a build gate (a quality gate that can block the build), and even when people forget, the script does not.


24.1.2 The Cascade of Three Checks

Verification is not one lump; it is stages. Run the cheapest check first to filter out the obvious violations, and pass only what survives to the next stage. Run the expensive checks on every input and the whole thing gets so slow that nobody runs it. Below is the verification flow I operate.

flowchart TD
    A[Doc or sheet saved] --> B{source_map audit}
    B -- source mapping missing --> B1[FAIL: traces of manual editing
_source_map.tsv update required] B -- pass --> C{wikilink integrity} C -- broken link found --> C1[wikilink_apply.py
healing attempt] C1 -- auto-heal possible --> C C1 -- unhealable --> C2[FAIL: broken reference report] C -- pass --> D{stale detection} D -- older than its target --> D1[WARN: added to review queue] D -- pass --> E[integrity_check final] E -- P0 violation --> E1[BLOCK: build gate blocked] E -- pass --> F[GREEN: commit allowed] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; classDef fail fill:#fee2e2,stroke:#dc2626,color:#7f1d1d; class B,C,C1,D,E code; class A data; class F pass; class B1,C2,D1,E1 fail;

The point of this cascade is that the earlier the failure, the cheaper it is. The source_map audit is a TSV line comparison and finishes in milliseconds. The final integrity_check, by contrast, loads the entire data sheet and checks foreign-key (FK) relationships, which takes several seconds. Put the cheap checks up front and the obvious mistakes get cut there; the expensive check runs only on the few inputs that made it through.

It also matters that each stage outputs differently. The audit emits FAIL (evidence that an editor touched something by hand); wikilink emits FAIL after attempting auto-healing; stale emits WARN (not blocking, but flagged for re-review); integrity_check emits BLOCK (stops the build itself). The same "problem" has to trigger a different response depending on severity, or people can't tell signal from noise.


24.1.3 Stage One — The _source_map.tsv Audit

The first check to run is source consistency. My document generation pipeline records, in _source_map.tsv, which synthesized document (say, the body of a game design document, GDD) came from which source files. Each line nails down the lineage: this output section = a synthesis of these source files.

This becomes a verification tool because hand-editing the output breaks the mapping. If someone directly edits an auto-generated GDD section, that section is no longer a faithful synthesis of its source files. The audit script compares the hash of each output section against a hash re-synthesized from the sources, and emits FAIL on a mismatch. That is where the rule "manual edits fail the audit" comes from.

This is not about preventing human edits — it is about making edits explicit. If an output needs fixing, the signal says: either fix the source and regenerate, or formally detach that section from the mapping (declare the separation) — one or the other. Making quiet edits loud is the audit's job.


24.1.4 Stage Two — Wikilink Integrity and Self-Healing

Pass the audit and you move on to the link check. My documents connect nodes with Obsidian-style wikilinks: [[target]]. wikilink_apply.py does two jobs — it resolves wikilinks to actual paths and applies them, and it heals broken links where it can.

The healable case is clear-cut: the target node was only renamed and still exists in the same place. A rename like the earlier 재료_목재_A재료_목재_상, as long as the alias map has been updated, gets auto-corrected from the old name to the new one by the script. If the target was deleted outright or can't be traced to wherever it went, the script gives up on healing and reports a broken reference.

There is one design judgment here. Overly aggressive auto-healing is dangerous. If the script hunts for a "similar name" and stitches links together on its own, links get attached to nodes with different meanings — a worse accident than the one you started with. So the healing in wikilink_apply.py is conservative — it auto-corrects only renames with an explicit alias mapping, and hands anything that requires guesswork over to a human. The virtue of automation lies in restraint: automate only what is certain, and honestly punt the ambiguous to people.


24.1.5 Stage Three — Stale Detection

Even when a link is alive, the reference can be stale. Document A references data sheet B; if B was updated later than A, then A's description may now disagree with the current B. The link itself is fine — the target it points at exists. But the content has rotted.

Stale detection compares the modification times (or content-hash versions) on both sides of a reference. If the referencing side is older than its target, it raises a WARN and puts that node in the review queue. It is WARN rather than BLOCK because an update doesn't always mean a content conflict. If the update fixed a single typo, the reference is fine. So stale is not a "blocker" — it is a "look at this" marker.

Here is how this stage would catch the broken-link incident from the opening. If the 재료_목재 sheet had been updated later than the document, a stale WARN would have fired before any auto-healing. In other words, the three checks are overlapping safety nets. What one net misses, the next net catches. That is why the cascade catches accidents no single check can.


24.1.6 Worked Transcript — Having Claude Build the Check Script

Writing all the verification logic from scratch wears you out. I describe the check rules in natural language and have the AI draft the executable script. Below is the actual session that produced the stale detection script. The output is not polished; the failures and rejections are left as they happened.

[Full Prompt]

The markdown files in my docs folder are connected to each other with wikilinks [[target]]. Write me a Python script that checks whether each document was updated longer ago than the target files it references (= stale). Conditions: - File modification time is based on the last git commit time (filesystem mtime can't be trusted — it gets refreshed on checkout) - Broken links (target missing) are excluded from the stale check and reported separately - Output is TSV: source<TAB>target<TAB>status (status = OK / STALE / BROKEN) - WARN is non-blocking; keep the exit code at 0 (must not block the build)

[Claude's Raw Output — First Pass]

The script below gets each file's last commit time with git log -1 --format=%ct <file>, extracts wikilinks with the regex \[\[([^\]]+)\]\], and compares.

import subprocess, re, sys, pathlib

def git_mtime(path):
    out = subprocess.run(
        ["git", "log", "-1", "--format=%ct", str(path)],
        capture_output=True, text=True)
    return int(out.stdout.strip())

def resolve(link, root):
    # [[target]] -> root/target.md
    return root / f"{link}.md"

root = pathlib.Path(sys.argv[1])
for md in root.glob("*.md"):
    src_t = git_mtime(md)
    for link in re.findall(r"\[\[([^\]]+)\]\]", md.read_text(encoding="utf-8")):
        tgt = resolve(link, root)
        if not tgt.exists():
            print(f"{md.name}\t{link}\tBROKEN")
            continue
        tgt_t = git_mtime(tgt)
        status = "STALE" if src_t < tgt_t else "OK"
        print(f"{md.name}\t{link}\t{status}")

This prints in the TSV format you requested.

[Human Verification / Rejection]

I stopped here. Three defects stood out.

First, if git_mtime returns an empty string, int("") blows up. That is exactly what happens with a new file that has never been committed. Second, glob("*.md") can't see subfolders. My documents are scattered under part*/. Third — and most important — wikilinks sometimes carry a pipe alias, as in [[target|display name]], and the first-pass code mistakes the whole target|display name for a filename. That flags every healthy link as BROKEN. Used as is, it is a false-alarm bomb.

I pointed out these three and re-requested.

[Re-request]

Fix three things: (1) for uncommitted files git_mtime is None — in that case withhold the STALE verdict and set status=UNTRACKED. (2) Use rglob instead of glob to include subfolders. (3) Handle the [[target|alias]] pipe — use only the part before | as the target. And keep the exit code at 0 even when there are BROKEN entries (the WARN non-blocking principle).

[Claude's Raw Output — Second Pass]

import subprocess, re, sys, pathlib

def git_mtime(path):
    out = subprocess.run(
        ["git", "log", "-1", "--format=%ct", str(path)],
        capture_output=True, text=True)
    s = out.stdout.strip()
    return int(s) if s else None   # None if no commit history

def target_name(link):
    return link.split("|", 1)[0].strip()   # [[target|alias]] -> target

root = pathlib.Path(sys.argv[1])
for md in root.rglob("*.md"):
    src_t = git_mtime(md)
    text = md.read_text(encoding="utf-8")
    for raw in re.findall(r"\[\[([^\]]+)\]\]", text):
        name = target_name(raw)
        hits = list(root.rglob(f"{name}.md"))
        if not hits:
            print(f"{md.name}\t{name}\tBROKEN")
            continue
        tgt_t = git_mtime(hits[0])
        if src_t is None or tgt_t is None:
            status = "UNTRACKED"
        elif src_t < tgt_t:
            status = "STALE"
        else:
            status = "OK"
        print(f"{md.name}\t{name}\t{status}")

sys.exit(0)   # WARN non-blocking: always 0

I accepted this second pass. One decision, though, remained with the human — when rglob(f"{name}.md") finds files with the same name in multiple folders, only hits[0] is used. That is a latent ambiguity, but under my document naming rules filenames are globally unique, so there are no collisions in practice. The AI never raised this assumption; a human consciously accepted it and left it in a comment. Even when automation wrote the code, the assumptions that code leans on are a human's responsibility.


24.1.7 Tying Check Results to the Build Gate

A script nobody runs is useless. The last piece of verification design is making it impossible not to run. I tie the three stages to a pre-commit hook or the build pipeline. An audit FAIL and an integrity_check P0 violation are BLOCK, stopping the commit or build; wikilink BROKEN and stale are WARN, allowed through but leaving a report.

This BLOCK/WARN split decides whether the system survives. Make everything BLOCK, and a single trivial stale blocks a commit — and people start routing around the verification itself. A verification that gets bypassed is a verification that doesn't exist. Leave everything at WARN, and the data integrity violations that genuinely must be stopped sail right through. The boundary between what you block and what you merely flag is the real design point of a verification system.


24.1.8 Measurement — Before and After Turning On Code Verification

These are directions observed on Project A at MMORPG studio A, where I run this system, at a scale of roughly 90 documents. Some of the absolute figures are the author's estimates (unverified); what carries meaning is the trend.

Item Manual-Review Era Code Verification Cascade
When broken references are found After player/QA reports Before commit (direction: post-incident → pre-incident)
Time for one consistency pass Several hours (author's estimate) Tens of seconds (measured from the script)
Stale dormancy Lurks for weeks WARN at the next commit
Bad auto-healing incidents N/A Held at 0 by conservative healing

Rather than taking the numbers at face value, I recommend trusting only the direction: the moment of discovery moved from after the fact to before it. The real value of a verification system lies less in the time saved than in the positional shift — accidents get caught before they reach players.


24.1.9 Common Failures

Pattern Prescription
Making every violation BLOCK, so people bypass verification Split BLOCK/WARN; block only data integrity P0
Auto-healing aggressively, all the way into guesswork Automate only explicit alias renames; ambiguity goes to a human
Watching only broken links and ignoring stale Detect aged references separately by comparing modification times
Quietly allowing manual edits to outputs Make edits visible as FAIL with the source_map audit
Having the scripts but not tying them to hooks Wire up pre-commit and the build gate, so they can't not run

Try It Yourself — A Minimal Verification Cascade

setup. Put your docs folder under git (the basis for commit-time comparison). Standardize wikilinks to the [[target]] or [[target|alias]] notation.

prompt. Give the AI the full prompt from the transcript above as is, but never use the first output as is. Without fail, verify these three — (1) handling of uncommitted files, (2) subfolder traversal, (3) pipe alias parsing — then reject and re-request. These are the spots the AI almost always misses on the first pass.

verify. Run the script and collect the TSV. Hand-check a sample of 5 BROKEN rows to confirm they are genuinely broken links. If false BROKENs show up, the alias or subfolder parsing is incomplete. Once it checks out, tie it to a pre-commit hook and branch the exit code: WARN (STALE/BROKEN) passes, BLOCK (data integrity P0) blocks.

Solo Scale-Down. If you are writing a small GDD alone, the full cascade is overkill. Take just the stale detection stage. Even comparing only whether a document is older than its data sheet by git time catches most of the "I thought I fixed it, but I didn't" accidents. Add auto-healing and the source_map audit once your documents pass 30 and you can no longer chase them by hand.


Key Takeaways

24.2 Mermaid Diagram Automation — Letting Documents Draw Their Own Diagrams

Three days into the job, a new designer asked me, "Is there a diagram somewhere that shows what order these systems affect each other in?" I hesitated. There was a diagram. A photo of something someone had drawn on a whiteboard six months earlier sat somewhere on the wiki. But in that picture, two systems that no longer exist were still alive, and three core loops added since were missing. In the end I answered, "Don't trust the diagram — read the docs." It was an embarrassing answer. The moment a diagram disagrees with the documents, it stops being information and becomes misinformation.

Let me state this chapter's conclusion up front. A diagram drawn by a human always rots within a month or two. So take the act of drawing diagrams out of human hands, and make the document structure itself spit out its own pictures. This chapter shows that process as one record of real work. I include, in full, a worked transcript that takes a document as input and generates Mermaid code, and the diagrams it produced are actually rendered on this very page. A chapter explaining a technique proves itself with that technique's own output.


24.2.1 Why Mermaid, of All Things

There are plenty of diagram tools: draw.io, Figma, Visio, even whiteboard photos. They all share one trap: the deliverable is an image file. An image can't be tracked line by line in git, can't be generated or edited directly by an LLM that works on text, and can't sit inside a Markdown document as code. From an operations standpoint, the first one is the most fatal. A picture where nobody can trace who changed what, when, and why becomes, over time, a relic nobody takes responsibility for.

Mermaid solves all three at once. You write the diagram as text, and the viewer handles the rendering. Because it's text, git diff catches even a single added node. Because it's text, an LLM can read and write it. Because it's text, it drops straight into a Markdown code block. The body of this very chapter is the proof. The diagrams coming up below this sentence are all text blocks inside Markdown, rendered into pictures during the book's build.

One misunderstanding needs heading off, though. There is no need at all to turn every operational artifact into a diagram. Listing items is faster with bullets; comparing numbers is faster with a table. Mermaid wins in exactly three places: relationships (what connects to what), flow (what comes after what), and sequence (who sends what to whom, and when). Force a diagram anywhere else and you only add cognitive load.


24.2.2 The Spine: One Session That Pulled a Diagram from Document Structure

This is where the chapter's backbone begins. Instead of abstract explanation, I show the entire process of turning one real chunk of a document into Mermaid, start to finish. The input is a Markdown fragment from Project A's ops documents that records the system dependency structure (below is a real excerpt, anonymized).

# System dependency memo (ops doc excerpt, anonymized)

- combat_core depends on stat_engine
- skill_runtime depends on combat_core
- skill_runtime depends on vfx_pool
- quest_director depends on skill_runtime
- quest_director depends on dialog_graph
- economy_loop subscribes to quest_director's reward hook
- economy_loop reads stat_engine's derived stats

Transcribing this into a diagram by hand means seven nodes and seven arrows. You can draw it once. The problem comes next week, when a mail_box system gets added and dialog_graph splits in two. From that moment on, the hand-drawn picture starts lying. So I hand this conversion to an LLM, not a person.

Step 1 — The Full Prompt

Below is the prompt I actually sent, included as is without polishing a single character.

Convert the system dependency memo below into a Mermaid graph (top to bottom, graph TB). Rules:
1. Only the systems in the memo become nodes; no adding new systems.
2. "A depends on B" becomes A --> B.
3. For weak couplings like "subscribes" and "reads", use a dotted -.-> with a relation name attached.
4. Use the English names in the memo as node ids, verbatim; do not attach Korean labels.
Code block only; no explanations or comments.

[Memo]
- combat_core depends on stat_engine
- skill_runtime depends on combat_core
- skill_runtime depends on vfx_pool
- quest_director depends on skill_runtime
- quest_director depends on dialog_graph
- economy_loop subscribes to quest_director's reward hook
- economy_loop reads stat_engine's derived stats

Rules 1 and 4 are the heart of it. The most common accident when you hand diagramming to an LLM is the "helpful hallucination": it adds nodes you never asked for ("most games have a login system, so I included one") and attaches Korean labels on its own, breaking the match between IDs and full-text search. The rules shut that helpfulness down.

Step 2 — Claude's Raw Output (No Summary)

What follows is the model's first response, reproduced as is, unrefined. There was one small problem mixed into it, and I include it, problem and all.



graph TB
    skill_runtime --> combat_core
    combat_core --> stat_engine
    skill_runtime --> vfx_pool
    quest_director --> skill_runtime
    quest_director --> dialog_graph
    economy_loop -.보상 훅 구독.-> quest_director
    economy_loop -.파생 스탯 읽기.-> stat_engine

The syntax is fine and there are exactly seven nodes. It didn't invent any new systems. Rules 1 and 4 held. But one thing bothered me. The dotted-edge labels came back in Korean (보상 훅 구독, "subscribes to the reward hook," and 파생 스탯 읽기, "reads the derived stats"). This isn't a rule violation. Rule 3 said to attach relation names, and the model complied faithfully. But I knew from experience that in Mermaid, a Korean label with spaces in it breaks some older renderers. The model had no way of knowing that.

Step 3 — Human Review and Rejection

This is the step most often skipped in automation, and the most important one. I didn't accept the output as is — I rejected it. Two reasons.

First, the Korean-with-spaces dotted-edge labels had to become English tokens to secure render compatibility. Second, weak couplings (dotted) and strong couplings (solid) were mixed in one picture with no color or style distinction, so it didn't read at a glance. I went back with these two points.

Step 4 — The Re-Request Prompt

Almost good. Fix only two things.

1. Change the dotted-arrow labels to a single English word each (no spaces). 
   "보상 훅 구독" -> subscribes, "파생 스탯 읽기" -> reads
   Reason: some renderers break on Korean-plus-space edge labels.
2. To visually distinguish dotted (weakly coupled) nodes from solid (strongly coupled) nodes,
   give nodes that have only weak couplings, like economy_loop, a pale gray style with classDef.
3. Leave everything else as is.

Step 5 — Raw Output for the Re-Request



graph TB
    skill_runtime --> combat_core
    combat_core --> stat_engine
    skill_runtime --> vfx_pool
    quest_director --> skill_runtime
    quest_director --> dialog_graph
    economy_loop -.subscribes.-> quest_director
    economy_loop -.reads.-> stat_engine

    classDef weak fill:#eee,stroke:#999,stroke-dasharray:4 3;
    class economy_loop weak;

This time I accepted it. The labels became single English tokens, and economy_loop alone drops into gray, so the information "this system is an edge system, tied in only by subscribing and reading rather than direct dependency" is carried by color. Had I drawn this by hand without ever touching a prompt, odds are I would never have thought of this classDef at all.

The Spine's Deliverable — Actually Rendered Right Here

The final output of the transcript above goes onto this book's page as the code block itself, not copied over by hand. The book's build draws it into a picture. This is "proving itself with its own technique," in the flesh.

graph TB
    skill_runtime --> combat_core
    combat_core --> stat_engine
    skill_runtime --> vfx_pool
    quest_director --> skill_runtime
    quest_director --> dialog_graph
    economy_loop -.subscribes.-> quest_director
    economy_loop -.reads.-> stat_engine

    classDef weak fill:#eee,stroke:#999,stroke-dasharray:4 3;
    class economy_loop weak;

One chunk of document excerpt, after five exchanges, became an operational asset that lives in git, that an LLM can update, and that renders on this page. When mail_box gets added next week, you write one line in the memo and throw the same prompt again. Nobody has to pick up a pen.


24.2.3 The Second Diagram: Drawing the Automation Pipeline Itself

If the previous diagram is "the result of the conversion," this one is "the process of the conversion." I turned the five-step worked procedure we just walked through into a flowchart. This diagram, too, was produced by tasking the LLM the same way, and it went through the same review. I include the result as is.

flowchart TD
    SRC[Ops doc excerpt] --> PROMPT[Write conversion prompt]
    PROMPT --> LLM[Claude raw output]
    LLM --> CHECK{Human review}
    CHECK -->|Reject: render compat / readability| REASK[Re-request prompt]
    REASK --> LLM
    CHECK -->|Approve| EMBED[Embed code block in Markdown]
    EMBED --> GIT[git commit and diff tracking]
    GIT -->|On doc change| SRC
    classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764;
    classDef human fill:#fde68a,stroke:#b45309,color:#000;
    classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b;
    class LLM ai;
    class PROMPT,CHECK,REASK human;
    class SRC,EMBED,GIT data;

This flowchart says one thing. What I want to emphasize — with a bold arrow, not a dotted one — is the diamond in the middle: Human review. Get drunk on the word "automation" and drop that node, and the helpful hallucination from Step 1 goes straight into your ops documents. Automation frees people from drawing; it does not free them from judgment. The loop's final arrow (On doc changeOps doc excerpt) is the key. Only with that feedback loop in place does the diagram stop being a one-off artifact and become an asset that grows with the document instead of aging apart from it.


24.2.4 The Conversion Script: A Deterministic Path That Runs Without an LLM

LLM conversion is flexible, but when the relationships already exist as structured data, there's no reason to call a model at all. For fixed-field data like Project A's decision cards, a small Python script is faster and more honest (hallucination is impossible by construction). Below is the core of the actual script that turns a list of decision cards into a decision-graph Mermaid.

# decision_graph_to_mermaid.py
# Decision cards (structured data) -> Mermaid graph. No LLM needed; deterministic.

def to_mermaid(decisions):
    lines = ["graph LR"]
    # 1) Declare nodes: id and title as is. Nothing gets invented.
    for d in decisions:
        safe_title = d.title.replace('"', "'")   # escape only the quotes
        lines.append(f'    {d.id}["{safe_title}"]')
    # 2) Edges: relation type becomes the arrow label.
    for d in decisions:
        for rel in d.relations:
            lines.append(f'    {d.id} -->|{rel.type}| {rel.target}')
    return "\n".join(lines)

The point is that it finishes in just two steps. Declare the nodes, connect the edges. A node absent from the input never appears in the output. Feed this script three decision cards and you get a graph like the one below.

graph LR
    D_A["Global cooldown 0.3s"] -->|superseded_by| D_B["Global cooldown 0.5s"]
    D_B -->|relates_to| D_C["Allow recovery-phase exception"]
    D_B -->|side_effect| D_D["Melee skill damage -5%"]
    classDef human fill:#fde68a,stroke:#b45309,color:#000;
    class D_A,D_B,D_C,D_D human;

One decision gets superseded by another (superseded_by), and even the side effect that branched off it (side_effect) shows up as a single arrow. Without reading dozens of lines of a text decision log, this one graph gets you the history of "why is the cooldown 0.5 seconds now" in under five minutes.

When to use the LLM and when to use a script? The criterion is simple. If the input is structured data (cards or sheets with fixed fields), use a script; if the input is free text (meeting notes, memos, conversations), use an LLM. Use an LLM on structured data and you take on hallucination risk for nothing; use a script on free text and the parsing rules grow without end.


24.2.5 Four Pitfalls and Their Remedies

These are mines I actually stepped on while running diagram automation.

First, the trap of growing too complex. Past fifty nodes, a picture no longer aids cognition — it obstructs it. The remedy is to cap a single view at twenty to thirty nodes, and once it grows beyond that, group regions with subgraph or split the diagram in two outright.

Second, the trap of updates going dead. It's tempting to think this only happens to hand-drawn pictures, but if you set up automation and then stop fixing the input document, it rots just the same. The remedy is the feedback loop from the earlier flowchart. Make the input document the single source of truth, and always regenerate the diagram from it.

Third, the trap of being too abstract. A picture at the level of "the systems are roughly tangled together like this" is pretty but useless. The remedy is to put real IDs (skill_runtime, D_B) on the nodes instead of abstract nouns. Full-text search and the diagram have to share the same identifiers before you can jump straight from picture to code.

Fourth, the trap of using LLM output as is, without review. As Step 3 of the spine showed, a model can follow every rule and still produce a label that breaks the renderer. The remedy is to never remove the human review node from the pipeline.


24.2.6 The Payoff — An Honest Account of the Change

The temptation to bring out numbers is real, but here I'll state direction only. The comparison below is the change I felt on the team I ran — an author's estimate (unverified), not a precise measurement.

The clearest change was how fast new hires came to understand the systems. Grasping "how these systems are tied together," which used to take the first several days on the job, shrank to about an hour in front of one auto-generated dependency graph. Meeting prep got lighter too. Someone used to redraw the picture by hand the night before a meeting; now converting the document once is the end of it. Above all, the question "can I trust this diagram?" — the one that surfaced whenever a diagram drifted from reality — has all but disappeared. The input document is the picture, so if the document is right, the picture is right.

In the other direction, to be honest: automation is not a cure-all. For early-stage idea sketches that haven't firmed up yet, a whiteboard is still faster. Automation shines after the structure has hardened to a degree.


Key Takeaways


Try It Yourself

setup. One Markdown repository holding your documents, plus a viewer that renders Mermaid (built into most Markdown viewers and git hosting), is enough. From the document you want to convert, pick one fragment that qualifies as "relationship, flow, or sequence" (for example, a system dependency memo).

prompt. Drop that fragment into the Step 1 prompt template from this chapter's spine and send it to the LLM. Always include the two rules: "do not add nodes that are not in the memo" and "use the original English names as IDs, verbatim." If the data is structured, convert it with a deterministic script like decision_graph_to_mermaid.py instead of an LLM.

verify. Paste the output code block into Markdown and actually render it. Check three things: (1) no nodes appeared that aren't in the input, (2) the edge labels draw without breaking, (3) the IDs used in your text and the diagram IDs match. If any one fails, reject with a re-request prompt and get a new version. Once it passes, commit to git — from now on, changes are tracked as diffs.

Solo Scale-Down

If you're a solo worker with no team and no scripts, shrink it down like this. In your note app, write the relationships between systems, tasks, and ideas as bullets in the form "A depends on B." Once a week, copy that whole list and throw a one-liner: "Convert this to a Mermaid graph TB; do not add nodes that are not in the list." Paste the returned code block at the top of the note. That's it. You never draw by hand, so there's no update burden, and as long as the input list stays alive, the picture is always current.

24.3 Wikilinks and Document Hierarchy — Links and Classification, the Two Entrances to Search

Links (wikilinks) and classification (hierarchy) are two entrances to the same problem. One answers "where does this decision lead"; the other answers "where does this document live."

On the morning of their second day, a newly joined designer asked me: "Is 0.5 seconds the right value for the combat global cooldown? Which document has the rationale?" I couldn't answer. I knew the decision was recorded somewhere, but I couldn't remember whether it was in the combat rulebook, the meeting notes, or a quarterly report. Three of us dug through the entire folder with grep. The same number turned up in six places, and we couldn't tell which was the "original decision" and which were "reference copies." We spent 40 minutes. What we finally found was a single line buried in meeting notes.

That evening I realized two things had been missing. First, there were no explicit links between documents. The same number lived in six places, but nowhere was there a thread saying "this one is quoted from over there." Second, the documents had no hierarchy to live in. Decision records were scattered across rulebooks, meeting notes, and reports, with no agreement that "decisions live here."

Those two things are the subject of this chapter. Wikilinks write the connections down as text; hierarchy promises the classification as folders. They look like separate techniques, but they are really two sides of one problem: search.


24.3.1 What Breaks When There Are No Links

With 30 documents, you keep it all in your head. Past 100, human memory stops working as the index. At that point you can rely on one of two things: sweep everything with grep (slow and imprecise), or follow the explicit links written inside the documents (fast and precise).

Why grep is imprecise is simple. Search for the string combat_global_cooldown_constant, and the document that decided the value and the documents that merely mention it come back looking identical. grep doesn't know which one is the original. But once we agree on the double-bracket notation [[combat_global_cooldown_constant]] inside documents, the signal "this intentionally references that atom" lives in the string itself. Narrow the search to the pattern \[\[combat_global_cooldown and accidental mentions drop out, leaving only intended references.

This one-line notation convention becomes an edge of a graph. When document A writes [[atom_X]], an A→X edge is created. With 200 documents each writing a few, the graph accumulates inside the text without anyone drawing it.

Below is a small piece of how our project's atoms, decisions, and documents are tied together by wikilinks. Node color indicates the kind; arrows indicate the direction of reference.

[[CombatFormula_v3]] [[Meeting_W21]] [[combat_global_ cooldown_constant]] [[D2026_Q2_017]] [[D2026_Q2_018]] Document atom Decision

What this small piece shows is that the answer to the new designer's question was already in the graph. Follow the arrows coming into the combat_global_cooldown_constant atom backwards and you reach decision D2026_Q2_017. Not 40 minutes — one reverse reference.


24.3.2 The Notation Convention — Four Kinds, One Format

We limited what wikilinks can point to — exactly four kinds. Add more kinds and the format wobbles; when the format wobbles, grep becomes imprecise again.

All four kinds share the single format [[name]]. The name must be globally unique. If an atom name collides across two places, they merge into the same node of the graph — the accident where "combat's cooldown" and "the UI's cooldown" become one node. That is why the atom naming rules force a domain prefix (combat_, ui_).


24.3.3 wikilink_apply.py — Apply and Heal

The notation convention alone isn't enough. Hand-bracketing 200 documents is unrealistic, and even once done, everything breaks the moment an atom is renamed. So we run a script that does two jobs. First, apply — automatically convert known atom names appearing in body text into wikilinks. Second, heal — find links that were renamed or broken, then update and report them.

The core of wikilink_apply.py looks like this.

# wikilink_apply.py — applies [[wikilink]] to atom names in body text and heals broken links
import re
from pathlib import Path

WIKILINK = re.compile(r"\[\[([A-Za-z0-9_]+)\]\]")
# Match only atom names appearing bare, not already linked (no [[ in front)
BARE_NAME = lambda name: re.compile(rf"(?<!\[\[)(?<![A-Za-z0-9_])({re.escape(name)})(?![A-Za-z0-9_])(?!\]\])")

def load_known_atoms(registry: Path) -> set[str]:
    # _atom_registry.tsv: first column is the currently valid atom name
    return {ln.split("\t")[0].strip()
            for ln in registry.read_text(encoding="utf-8").splitlines()
            if ln.strip() and not ln.startswith("#")}

def apply_links(text: str, known: set[str]) -> tuple[str, int]:
    applied = 0
    for name in sorted(known, key=len, reverse=True):  # longest names first: prevents partial-match corruption
        text, n = BARE_NAME(name).subn(rf"[[{name}]]", text)
        applied += n
    return text, applied

def heal_links(text: str, known: set[str], aliases: dict[str, str]) -> tuple[str, list[str]]:
    dead = []
    def repl(m):
        ref = m.group(1)
        if ref in known:
            return m.group(0)              # alive → keep as is
        if ref in aliases:                  # renamed atom → heal to the new name
            return f"[[{aliases[ref]}]]"
        dead.append(ref)                    # genuinely dead link → report
        return m.group(0)
    return WIKILINK.sub(repl, text), dead

Two design choices here are the backbone of the script.

First, apply_links replaces longer names first. With two atoms combat_cooldown and combat_cooldown_global, replacing the shorter one first corrupts the front of the longer one. A single line of descending-length sorting prevents that accident. This is the part I missed when I first wrote the script, and I only added it after the broken result [[combat_cooldown]]_global actually showed up.

Second, heal_links heals through a rename dictionary (aliases). When an atom name changes from combat_gcd to combat_global_cooldown_constant, the old name is automatically replaced with the new one, and only when the dictionary doesn't have it either is it reported as a dead link. Instead of hand-fixing 200 documents on every rename, we add one alias line.


24.3.4 Worked Transcript — Handing Wikilink Augmentation to Claude

Automatic apply only links "atom names already known." It can't catch sentences that spell things out without using the atom name ("The combat global cooldown is 0.5 seconds"). Turning these natural-language references into wikilink candidates is something an LLM does faster than a person. Below is the actual exchange, in full. The output is not summarized, and the part I rejected is left in as is.

My prompt (in full):

I'll give you one paragraph from the combat rulebook and the list of registered
atoms. Find the natural-language expressions in the text that point to these
atoms and propose [[wikilink]] candidates as a table. Only when you're certain;
if it's ambiguous, just mark it "ambiguous". Don't rewrite the text — only say
which phrase maps to which link. Pull unregistered concepts out separately as
"new atom candidates".

[Registered atoms]
combat_global_cooldown_constant
combat_skill_cast_interrupt_rule
combat_resource_regen_rate

[Text]
The combat global cooldown is 0.5 seconds and applies to all active skills.
Getting hit while casting interrupts the cast. Healing-type skills are the exception.
Resource regen returns to normal speed only while out of combat.

Claude's output (raw):

Proposed wikilinks:

| Text phrase                                      | Candidate link                          | Confidence |
|--------------------------------------------------|----------------------------------------|------|
| "The combat global cooldown is 0.5 seconds"      | [[combat_global_cooldown_constant]]    | High |
| "Getting hit while casting interrupts the cast"  | [[combat_skill_cast_interrupt_rule]]   | High |
| "Resource regen ... only while out of combat"    | [[combat_resource_regen_rate]]         | Medium |

New atom candidates:
- "Healing-type skills are the exception" → This exception rule is not spelled
  out in any of the registered atoms. It looks like it needs a separate atom,
  something like combat_skill_cast_interrupt_heal_exception. (Ambiguous: it
  could be a sub-clause of interrupt_rule, so I can't judge whether it should
  be a standalone atom.)

Note: for the "resource regen" row, I haven't seen the atom definition, so I
can't tell whether combat_resource_regen_rate refers to the 'regen rate' or to
the 'out-of-combat condition' — hence medium confidence.

My verification and rejection:

I accepted two of the three rows. The global cooldown and the cast interrupt matched the atom definitions exactly. On the third "resource regen" row, Claude was right to lower its own confidence to "medium" — the combat_resource_regen_rate atom defines only the rate value, and the condition "only while out of combat" belonged to a different atom. Linking it as proposed would wrongly attach a "condition" to a "rate" atom. I rejected it.

The new-atom-candidate call was accurate. The "healing exception" really had no atom anywhere. The part Claude flagged as ambiguous ("a sub-clause of interrupt_rule, or a standalone atom?"), however, was a human's call to make, and I decided to split it out as a standalone atom.

Re-request:

Don't link the "resource regen" row. Instead, [[combat_resource_regen_rate]]
covers the 'rate' only, and the 'out-of-combat condition' is split into a new
atom. Write a one-line definition for each of the two atoms. Also write a
one-line definition for the healing exception as a standalone atom.

What the LLM did in this round-trip was not "creating candidates from zero" but "picking candidates out." The point is that there was a clearly marked place for a human to refuse. Had this been auto-published, one wrong link would have stayed in the graph permanently.


24.3.5 Lint — Stopping Broken Links at Build Time

Links break over time. Atoms get retired, names change, typos creep in. So we run a wikilink lint on every build. The checks and their handling:

Keeping dead links a warning rather than a block is deliberate. In the middle of renaming an atom, dead links appear briefly, and failing the build over that would stop work. Instead, the lint makes you check the heal dictionary first. Format violations and name collisions are blocked immediately — those two corrupt the entire graph.

This lint is self-evidencing. The links that wikilink_apply.py creates are checked by the same system's lint, and the result is recorded as yet another atom decision. A tool verifying its own output against its own standards — that loop is the basic skeleton of operations.


24.3.6 Classification — The Hierarchy Where Documents Live

That covers links. Now, classification. If wikilinks answer "where does this decision lead," hierarchy answers "where does this document live." Without both, the new designer's 40-minute search repeats.

Our document folder has four layers. The layers share the same skeleton as the Layer unification in our information architecture — vision, systems, content, and meta each get one layer.

docs/
├── L0_vision/              Vision (5 docs or fewer, rarely changes)
├── L1_systems/             Per-domain rulebooks
│   ├── combat/
│   ├── narrative/
│   └── ui/
├── L2_content/             Individual content
│   ├── characters/
│   └── quests/
└── L4_meta/                Ops, decisions, meetings, atoms
    ├── decisions/
    ├── meetings/
    ├── reports/
    └── atoms/

L3 is empty because data sheets and the DB occupy that spot (tables, not documents). The decision the new designer was hunting for lives in L4_meta/decisions/ — with just that one promise in place, the 40-minute search would have ended with one sentence: "decisions live there."

For the hierarchy to work as a search entrance, five things must hold together. Drop any one of them and the classification collapses.

  1. Classify by meaning, never by time. combat/ and narrative/ stay searchable; nobody opens 2026-Q1/ or 2026-Q2/ six months later. Git records time, so there is no reason to split folders by it again.
  2. Depth 4 or less. L1_systems/combat/skills/active/single_target/attack.md is five levels. Once a path overflows one screen, people can't hold the location in their head.
  3. Filename prefixes. Put the kind into the filename with spec_, report_, decision_, char_. You can see the kind without looking at the folder.
  4. A README in every folder. Each folder's README states its definition, contents, and naming rules. It is a newcomer's first entrance.
  5. _-prefixed meta folders. _archive/, _TEMPLATES/, _NAMING/ sort to the top automatically and never mix with the real content.

A document doesn't stay in one place. While being written it lives in its home folder with status: draft; once activated it becomes status: active; when retired it is not deleted but moved to _archive/ and tagged status: deprecated. Never deleting deprecated material is an iron rule. Six months later, when someone asks "why did we reverse that decision?", the answer exists only inside the deprecated material. Delete it, and there is no way to recover the decision's rationale after the fact.

Large changes are not left to git alone; they are recorded as a change_log in the frontmatter.

---
title: combat_global_cooldown_rule
version: v3
last_modifier: teammate_a
change_log:
  - v1 (2025): initial draft
  - v2 (2025): cooldown 0.3 → 0.5  ([[D2026_Q2_017]])
  - v3 (2026): added healing exception  ([[D2026_Q2_018]])
---

Notice that the decision IDs in the change_log are written as wikilinks. This is where links and classification meet. The document lives in one place inside the hierarchy (classification), while its change history leads into the decision graph (links). One frontmatter opens both entrances at once.


24.3.7 Once a Quarter, the Cleanup Cycle

Left alone, a hierarchy rots. Empty folders appear, six-month-old drafts pile up, depth creeps upward. So we clean up once a quarter. Delete empty folders; decide activate-or-retire for drafts older than six months; flatten anything at depth 5 or more; write or retire READMEs for folders missing one; and when _archive passes half the total, compress and preserve it. Without this cycle, the hierarchy fills with noise until signal and noise become indistinguishable.

The whole flow in one picture: a document comes in, gets linked, classified, verified, and eventually retired — one loop.

flowchart TD
    A[New document written
status: draft] --> B[wikilink_apply.py
auto-links atom names] B --> C[LLM augmentation
natural-language reference candidates] C --> D{Human review} D -->|accept| E[Hierarchy placement
L0~L4 + prefix] D -->|reject| C E --> F[Build lint
dead/collision/cycle checks] F -->|pass| G[status: active] F -->|dead link| H[Check heal dictionary] H --> F G --> I[Quarterly cleanup cycle] I -->|retire| J[_archive/
status: deprecated] I -->|keep| G classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d; classDef fail fill:#fee2e2,stroke:#dc2626,color:#7f1d1d; class B,F,H code; class C ai; class D,I human; class A data; class G pass; class J fail;

In this loop, links (B·C·D·F) and classification (E·I·J) take turns. They don't run separately; they interlock within the life of a single document.


24.3.8 Effects — What Changed and How

The numbers here are directional, comparing before and after adoption on my project. They are not precise measurements but the size of the difference I felt doing the same work in the two environments (author's observation, not precisely measured).

Before links and hierarchy were in place, a new designer's decision-tracing question took up to an hour or two, like the 40-minute case in the opening. After, it's one atom reverse reference — a matter of minutes. Document search dropped from 5–10 minutes to around 30 seconds, the result of the hierarchy's meaning-based classification and filename prefixes working together. Accidents caused by wrong citations (the kind where an already-retired value is mistaken for the current one) fell from several per quarter to one or two — with wikilinks stating "this references that atom," copied values and original values stopped being confusable.

The biggest change was newcomer onboarding. Without hierarchy, learning which folder held what took days; without links, there was no way to grasp how the systems intertwined. With both in place, new joiners learned locations from the folder READMEs and explored cross-system relationships on their own by following the wikilink graph. "Things you can only learn by asking" turned into "things you can see by following."

This effect appears only when both entrances exist together. Links without classification: the graph is there, but you don't know where documents live. Classification without links: the folders are tidy, but you don't know where decisions lead.


24.3.9 Common Failures and Remedies

On the links side, the most common failure is noise links. If you bracket every noun because wikilinks seem great, the graph fills with meaningless edges and visualization tools become useless. The principle: keep only links for which you can ask and answer "what is the relationship between this document and that one." Next is auto-publishing — commit LLM-made links without human review, and a wrong connection like the "resource regen" row in the worked transcript stays forever. Apply is automatic; publishing is human.

Failures on the classification side are mostly violations of the five principles: time-based folders, depth 5 or more, lawless filenames, missing READMEs. And the hardest one to undo — deleting deprecated material. The rationale behind a deleted decision cannot be recreated. The one line that moves it to _archive protects the learning material you'll need six months later.


Key Takeaways


Try It Yourself — A Minimal Wikilink + Hierarchy Adoption

setup. Create four folders in your docs folder — L0_vision/ L1_systems/ L2_content/ L4_meta/ — and put a one-line README in each. Collect your atom names into a single _atom_registry.tsv file (first column = atom name).

prompt. Give the LLM one paragraph of text and the registered atom list, and ask: "Find the natural-language expressions in the text that point to these atoms and propose [[wikilink]] candidates as a table. Propose only when certain; if ambiguous, just mark it 'ambiguous.' Do not rewrite the text. Separate unregistered concepts as 'new atom candidates.'"

verify. Check every proposed link against the atom definition. Accept only when what the atom refers to and what the text refers to are exactly the same; reject when a condition, attribute, or exception diverges. After accepting, run grep "\[\[name\]\]" to confirm the links actually landed and there are no dead links.

Solo Scale-Down. To start with no script and no lint, two rules are enough. (1) Decisions go in one folder, decisions/, as decision_*.md — always. (2) When another document mentions a decision, write it as [[decision_id]]. Keep just these two rules, and you can answer "where is that decision?" with a single grep "\[\[decision_". Bringing in tooling after your documents pass 100 is not too late.

24.4 Source Tracking and Data Lineage

The moment you start doubting a piece of data always arrives too late. Only after a wrong number has gone into the live build do you ask, "Where did this come from?"


On the Friday evening right before the alpha build, team member B came to my desk, laptop in hand, the combat balance spreadsheet open on the screen. "Director, the sheet says the boss's phase 1 HP is 48,000, but the value that went into the build is 52,000. Which one is right?"

I don't know. To be precise — at that moment, nobody knows. The 52,000 could be the latest value, reflecting a decision from a meeting a few days back, or it could be an unverified value someone parked there temporarily. The 48,000 could be the agreed value from before that meeting. Both numbers are plausible. Plausibility is not evidence.

To answer the question, you have to trace back to the source. Which meeting decided it, what were that meeting's inputs, who copied it into the sheet. But if that chain of tracing exists only in people's memories, the answer becomes "I'll ask team member A tomorrow." Six months into live operations (live ops), unresolved questions like that pile up into a mountain. Data lineage — the genealogy of your data — is the infrastructure that keeps that mountain from forming.

The core principle is a single one: never record sources by hand. Source records that people backfill after the fact don't survive a month. Only sources recorded automatically, at the moment the data is created, survive.


24.4.1 The Five Costs of Data with Broken Sources

Recording one line of _source_map.tsv automatically costs a few milliseconds. The cost you pay when that line is missing spreads in five directions.

Broken source (source not recorded) Unverifiable "Where is this number from?" Missed updates source updated → derivatives stale Legal exposure basis for external assets lost Slow incident diagnosis bad values cannot be traced back Handover loss no answer to "why this decision?"

The trap is that none of the five costs is visible at the moment the data is created. Every invoice arrives weeks later, months later, after the people have changed. That is why sources can never be a "we'll clean it up later" item. They must be recorded at the moment of creation.


24.4.2 _source_map.tsv — The Standard Skeleton of Source Mapping

Project A runs exactly one source-mapping file: _source_map.tsv. The reason it is tab-separated text is simple. A human can read one line at a glance, a script parses it with a single split('\t'), and git diff shows a one-line change cleanly. CSV breaks when commas appear inside the content, and JSON makes a single line hard for a human to read.

asset_id    source_type source  created creator notes
spec_combat_v3  internal    mtg_battle_2026-04-18   2026-04-18  teammate_a  based on decision_D2026_Q2_017
data_boss_hp_v3 internal    decision_D2026_Q2_017   2026-04-18  teammate_b  phase 1 48000 confirmed
asset_K_001_concept internal_ai_assisted    imagegen + teammate_b cleanup   2026-04-20  teammate_b  legal_review done
data_user_voice_W21 external_aggregated forum + community + sns 2026-05-25  auto_collect    output of the 13.1 pipeline
ref_visual_tone_a   external_reference  refgame (2024)  2026-04-15  teammate_c  visual tone reference, no direct borrowing

The six columns have clear roles. asset_id is the data's unique key; source_type is the classification (covered below); source is where it came from — a meeting ID, a decision ID, a collection pipeline, an external work; created/creator are when and by whom; notes is one line of context for humans to read.

Now look at the second and third rows again, and the answer to team member B's question from the previous section appears. The source of data_boss_hp_v3 is decision_D2026_Q2_017, and its notes read "phase 1 48000 confirmed." The build's 52,000 is nowhere in this lineage. In other words, 52,000 is an unverified temporary value, and the correct answer is 48,000. The question closes in one or two minutes — without calling on anyone's memory, and without ruining a Friday evening.

There is one more rule attached to this file, though. If a human edits _source_map.tsv by hand, the audit in integrity_check reports FAIL. The reason comes in a later section — sources must be recorded only automatically.


24.4.3 The Five source_type Categories — Classification Is the Processing Rule

The reason sources are classified into five types is not tidiness for its own sake. Each source_type carries its own operating rules.

internal meetings, proposals, decisions → tracking only internal_ai_assisted AI-generated + human cleanup → cite source external_aggregated user measurement → note date and sample external_reference third-party works → legal_review required self_measured Classification → processing rule mapping internal family: passes if traceable to a decision ID ai_assisted: notes must say which tool, which prompt aggregated: numbers uninterpretable without a collection date reference: empty legal_review → audit FAIL ← enforced self_measured: sims/KPIs; repro conditions in notes advised → source_type is not a label but a switch the checker reads and branches on

Take the external_reference row. If an asset drew on refgame as a visual tone reference, that asset must not go into a build without legal review. When the source_type is external_reference and the legal_review record is empty, the audit blocks it. This is the point where the label stops being just a label and becomes a switch the checker reads. When I say the five-type classification is the skeleton of operational trust, this enforcement is what I mean.


24.4.4 Automatic Recording — One Line Left at the Moment of Creation

Now the core. Sources must be recorded automatically at the moment data is created. Project A's source_tracker.py is hooked into asset creation.

# source_tracker.py
import time, getpass, csv
from pathlib import Path

SOURCE_MAP = Path("_source_map.tsv")
VALID_TYPES = {
    "internal", "internal_ai_assisted",
    "external_aggregated", "external_reference", "self_measured",
}

def track_source(asset_id: str, source_type: str, source: str, notes: str = ""):
    if source_type not in VALID_TYPES:
        raise ValueError(f"unknown source_type: {source_type}")
    if source_type == "external_reference" and "legal_review" not in notes:
        raise ValueError(f"{asset_id}: external_reference assets must be marked legal_review")

    record = [
        asset_id,
        source_type,
        source,
        time.strftime("%Y-%m-%d"),
        getpass.getuser(),
        notes,
    ]
    with SOURCE_MAP.open("a", encoding="utf-8", newline="") as f:
        csv.writer(f, delimiter="\t").writerow(record)

With this function hooked into the asset creation pipeline — when a sheet is exported, when a concept asset is registered, when user data is aggregated — one line of source is appended automatically. There is no step a human can forget. The burden of backfilling drops close to zero.

Auto-filling the creator column with getpass.getuser() is a small but decisive detail. Make people type their own names and you get blank cells. Have the system write it and there are none. And when an external_reference carries no legal_review, the record itself is refused — the channel through which an unsourced external asset could enter the system is closed off.

Here I want to pin down what using AI actually looked like. I got the first draft of this track_source function from Claude. What follows is one spine of that worked transcript.

My prompt (verbatim):

Write me a Python function that appends one source line to _source_map.tsv.
Columns: asset_id, source_type, source, created (today's date),
creator (current OS user), notes. Tab-separated.
Restrict source_type to five values: internal, internal_ai_assisted,
external_aggregated, external_reference, self_measured.

Claude's raw output (not summarized):

import csv, datetime, os

def track_source(asset_id, source_type, source, notes=""):
    record = [asset_id, source_type, source,
              datetime.date.today().isoformat(),
              os.getlogin(),          # ← current user
              notes]
    with open("_source_map.tsv", "a", newline="") as f:
        csv.writer(f, delimiter="\t").writerow(record)

If you want to add source_type validation, you can build a VALID_TYPES set and check it at the top of the function. Shall I add that for you?

My verification and rejection:

This output works, but in two places it did not fit our operation. First, os.getlogin(), when called from cron or CI, can return an empty string or raise an exception depending on the environment. Our export pipeline also runs on an unattended schedule. So I switched to getpass.getuser() — it consults environment variables and resolves the user more reliably. Second, Claude left source_type validation as an option — "shall I add that for you?" — but for us that is not an option, it is mandatory. Without validation, a typoed source_type gets in and the classification collapses.

My follow-up request:

Switch to getpass.getuser(). And make source_type validation mandatory,
baked into the function, not optional. Also, when the type is
external_reference and notes does not contain the string legal_review,
raise a ValueError. I want to block at the source any external asset
being recorded without legal review.

The result of this follow-up is the final source_tracker.py shown above. The point worth marking is not that Claude's first output was wrong, but that the rejection and follow-up were needed because I knew operational constraints the AI did not — the unattended schedule, the legal_review mandate. The AI delivers generally correct code fast; the human verifies whether it is correct in our environment. That verification point becomes a design decision of the source tracking system itself.


24.4.5 audit FAIL — The Integrity Check That Blocks Manual Edits

I said earlier that if a human hand-edits _source_map.tsv, integrity_check reports FAIL. How is that caught?

The principle is simple. Every time track_source appends a line, it serializes the line's key columns (asset_id, source_type, source, created, creator) into a hash and accumulates it in a separate .source_map.audit file. The audit check re-reads _source_map.tsv, recomputes the hashes the same way, and compares the two hash lists.

# the source_map audit part of integrity_check
def audit_source_map():
    fails = []
    rows = read_tsv(SOURCE_MAP)
    expected = read_lines(AUDIT_FILE)   # hashes accumulated at append time

    for i, row in enumerate(rows):
        h = row_hash(row["asset_id"], row["source_type"],
                     row["source"], row["created"], row["creator"])
        if i >= len(expected) or h != expected[i]:
            fails.append(f"L{i+1} {row['asset_id']}: suspected manual edit (hash mismatch)")

    if len(rows) != len(expected):
        fails.append(f"row count mismatch: tsv={len(rows)} audit={len(expected)}")
    return fails

Suppose someone hand-edits the source of data_boss_hp_v3 in the sheet to decision_D2026_Q2_099. The hash of that line no longer matches the original hash accumulated in the audit file, and the check prints this.

[FAIL] source_map audit
  L3 data_boss_hp_v3: suspected manual edit (hash mismatch)
  → Change bypassed track_source(). Sources must be recorded via the code path only.

Why does this enforcement matter? Allow hand edits, and sooner or later someone under pressure fills in a source that merely looks plausible. At that moment the lineage stops being the truth and degrades into a file of somebody's guesses. The audit FAIL gives teeth to the rule that sources go through the automated path only. The verification system of §24.1 bundles this audit with the other checks and runs it in CI.


24.4.6 Change Propagation — When the Original Changes, Wake the Derivatives

The real reason to record sources automatically is the reverse query: "Original X changed. What is affected?"

def find_derivatives(source_id: str):
    return [
        row for row in read_tsv(SOURCE_MAP)
        if row["source"] == source_id
    ]

# usage: decision_D2026_Q2_017 was overturned in a meeting
deps = find_derivatives("decision_D2026_Q2_017")
# → [spec_combat_v3, data_boss_hp_v3, ...]

Say decision_D2026_Q2_017 is overturned at the next meeting and the boss's phase 1 HP changes from 48,000 to 50,000. Call find_derivatives and every derived asset hanging off that decision comes back immediately — the combat spec document, the HP data sheet. Each asset's owner gets notified, and the incident of "an asset still pointing at an old decision" surviving into the build drops from several per quarter to nearly zero.

With hand-written sources, this reverse query never holds. If sources are free text, decision_D2026_Q2_017 gets written as "the Q2 017 decision" on one line and "the decision from Q2 meeting no. 17" on another, and matching breaks. Only with _source_map.tsv's standard format and track_source's automatic recording does change propagation actually work.


24.4.7 The Lineage Graph — A Data Genealogy in One Screen

Read line by line, _source_map.tsv is flat — but as one asset's source becomes another asset's source, the data's genealogy forms a chain. Unfold that chain on one screen and the trustworthiness of a decision's inputs becomes visible. This mermaid diagram is generated directly from _source_map.tsv by the automated diagram pipeline of §24.2 — the technique proving its own assets, so to speak.

graph LR
    Meeting["mtg_battle_2026-04-18
(meeting)"] --> Proposal["P2026_Q2_017
(proposal)"] Proposal --> Decision["D2026_Q2_017
(decision: 48000 confirmed)"] Decision --> Spec["spec_combat_v3
(combat spec)"] Decision --> Data["data_boss_hp_v3
(HP sheet)"] Data --> Build["build_2026-05-20
(alpha build)"] Build --> UserData["data_user_voice_W21
(user measurement)"] UserData -.input to the next decision.-> Decision classDef human fill:#fde68a,stroke:#b45309,color:#000; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; class Meeting,Decision human; class Proposal,Spec,Data,Build,UserData data;

A cycle forms naturally. The build produces user data, and the user data becomes the input to the next decision. Once that cycle is visible, "where did this number come from" becomes a path on the screen. Team member B's Friday question, on this graph, is one hop back along Data → Decision.


24.4.8 Measurement — The Effects of Running Lineage

On Project A I compared before and after adopting the lineage system. The time figures below are the author's estimates (unverified); read them for direction and ratio rather than absolute values. The counts are measured, tallied from the quarterly audit logs.

Item Without lineage With lineage Nature
Time to identify a data source 1–2 hours 1–2 minutes author's estimate (unverified)
Basis for verifying data trust a senior's memory immediate source lookup qualitative
Derivatives missed on source change 5–8 per quarter 0–1 measured from audit logs
External assets missing legal review possible 0 (recording enforced) measured from audit logs
Quarterly audit duration 1–2 days 2–3 hours author's estimate (unverified)

The hardest number is the "derivatives missed on source change" row. It can be counted because the audit log keeps the decision ID and the missed derivative assets verbatim. The time figures depend heavily on the measurement environment (team size, asset count), so I marked them as estimates. The direction is unambiguous — once sources are recorded automatically, tracing shifts from memory to lookup.


24.4.9 Common Failures and Remedies

Failure pattern Remedy
Sources backfilled by hand after the fact automatic recording at creation time via track_source
Source format differs from line to line _source_map.tsv tab standard + format enforcement
External assets missing legal review legal_review enforced in source_type validation
Hand-editing _source_map.tsv hash-comparison FAIL via the integrity_check audit
Derivatives left stale when the original changes find_derivatives reverse query + notifications
Genealogy explained only in prose one-screen visualization via auto-generated mermaid

What the six remedies share is that none of them depends on human diligence. Automatic recording, format enforcement, hash comparison, reverse queries — the system does all of it. Because the single reason source tracking collapses is that people forget.


24.4.10 Closing Part 24

Part 24 has been four threads of propping up operational trust with automation: chapter 1 gathered verification into one point with the verification system, chapter 2 drew the structure with automated mermaid generation, chapter 3 connected and layered the documents with wikilinks and the document hierarchy, and finally this chapter 4 sealed the trustworthiness of the data with sources and lineage.

The one sentence running through all four chapters is this: operational trust comes from the system's records, not from people's memories. Just as verification automatically asks "does this artifact conform to the rules," lineage automatically answers "where did this data come from." The crux of both is that they don't collapse even when people forget.

This operating know-how runs along the same grain as the book's Layer-unified design philosophy. Vision (assetization, trust) descends into systems (source rules), systems into data (_source_map.tsv), and data into build and QA (audit, automatic refresh) — one chain. That chain itself is lineage.


Key Takeaways


Try It Yourself (setup → prompt → verify)

setup. Create _source_map.tsv at the project root with a single header line (asset_id\tsource_type\tsource\tcreated\tcreator\tnotes), and put the source_tracker.py above next to it. Hook a track_source(...) call into the end of your asset export and registration scripts.

prompt. When you need an automatic source-recording function, ask Claude like this.

Write me a Python function that appends one line to _source_map.tsv
(tab-separated; columns: asset_id, source_type, source, created, creator, notes).
Restrict source_type to the five types, and when it is external_reference
with no legal_review in notes, raise a ValueError. Use getpass.getuser() for creator.

verify. Check two things yourself. (1) Call it with external_reference but leave notes empty, and confirm a ValueError is raised. (2) Change one character in the source column of _source_map.tsv with a text editor, run the source_map audit in integrity_check, and confirm it reports FAIL. If both are blocked, the source path is sealed.

Solo Scale-Down

If you work alone, one file — _source_map.tsv — and one function — track_source — are enough. Add the integrity audit, the reverse query, and the mermaid automation one at a time, once your assets pass a few dozen and the sources start to blur. The starting point is one habit: whenever you write down a number, leave one line of source in the same place, automatically.

Appendix

Epilogue — From the Chat Window to the Game Design Room

This book began in a single LLM chat window. It started with one line — "Clean up this design doc for me" — and from there, AI walked into the entire game design room. Where we stand now is not the end of that change, but the tying-off of one stage.


What Changed

Looking back from twenty-four years as a game designer, some things have clearly changed. Mass production has been absorbed into tools, meeting notes have become assets, and decisions have become traceable.

That said, the change has not reached every discipline at the same speed. Systems design and balance absorbed the tools quickly, while narrative and art direction often still sit at the conservative-application stage. Which discipline is faster is not the point. The next decision starts from accepting that change enters each discipline in a different place.

This does not mean designers now do less work. It means we do different work in the same hours. We moved from mass production to meaning, from organizing to deciding. Meaningful work does not automatically rush in to fill the space the tools cleared, so an awkward period follows for a while — one where you have to start by deciding, all over again, what not to do.


What Did Not Change

Some things did not change. Games are for people, a game's vision is decided by people, and the promise to make games that respect the player's time stands as it always has.

Tools are tools; direction belongs to people. But I intend to re-examine that one line every year. If I don't, then wherever the tools have grown strong enough, a default quietly settles in — a vague faith that "the decisions are still made by people." The moment "direction belongs to people" starts to sound like a safe promise is, in fact, the most dangerous moment of all.


What Comes Next

This book is a record of one moment in time. AI tools evolve fast, and a year from now parts of this book will likely be dated. But I believe the core patterns — Layer integration, decision tracking, verification gates, human review, team consensus — hold even as the tools change. If anything, the stronger the tools get, the faster any use that does not run on these patterns collapses.

Layer integration is not merely about unifying the language of collaboration across disciplines; it is about opening the road to procedural generation and automation ahead of time. As the stages widen — from the conservative application where a writer injects context one line at a time, to the progressive application where the system generates automatically from world state — human review just before the irreversible steps, like recording, capture, and live builds, serves as the last safety net. As LLMs get smarter, the value of this skeleton does not shrink. If anything, the weight of the decisions a human must review grows heavier.

If you take the patterns in this book and rework them for your own environment, then in effect you are writing this book's next edition. Depending on studio size, genre, development stage, and team makeup, some patterns will carry over as they are and some will need to be rebuilt. Telling apart what carries over from what needs rebuilding is itself the first meaningful piece of work.


Acknowledgments

Above all, my deep thanks go to the executives and the CEO of SCYBS Games for allowing this book to be published. Without their decision to open the way for workflows that grew naturally inside the company to be shared with the industry, this book would not exist. I thank every colleague who walked the road of game development with me — and sometimes apart from me — over twenty-four years; my partner, who has stayed by my side for more than twenty years; and Kkongji, the Persian cat still keeping me company at twenty-three. The same gratitude goes to Gomi, the Pomeranian who is no longer with us and will be nineteen forever.

I believe one experience of early success turning into a poisoned chalice is enough for a lifetime. So as not to forget that one time, I have spent close to twenty-four years honing my craft and learning again. This book, too, is something like a knot tied at one point along that learning.

Finally, my thanks to the AI tools that helped in writing this book. In the time it took to close this manuscript, Claude's new model Fable was released — that is how fast this landscape is changing, day by day.

If this book can add even one line to someone's next decision, that is enough.


Minsoo Lee, 2026

Appendix A. Detailed Inventory of the Company PC System

This appendix gathers the company PC environment I run at MMORPG studio A onto a single page — from hardware to tools, knowledge assets, and verification materials. Throughout the main text I mention "with this tool," "with this atom structure," "with this report"; this is where you can see, at a glance, the actual scale and combination in which those things exist. All real names and proper nouns have been anonymized, and since the figures shift with time, please read them as ratios and composition rather than absolute values.

There are two ways to read this appendix. One is to compare it with your own environment, item by item. Fill in the same cells — "Which engine do I use, which collaboration tools, in what form do my knowledge assets accumulate?" — and your blank cells reveal themselves. The other is to look at the balance of the composition. Having many tools does not make a good environment; what matters is whether the five axes — engine, design, art, collaboration, and AI — mesh together without blocking one another. Please look at the weave rather than at any single item.


A.1 System Overview

First, the hardware and operating system that form the foundation. Once you start using AI tools in earnest, you end up running diffusion models or STT locally, so memory and GPU headroom translate directly into working speed. Treat the specs below as something close to a lower bound — the baseline at which "things run without stalling."

A.1.1 Hardware

Item Spec
CPU Desktop workstation class
RAM 64GB or more
GPU For UE5 development (CUDA compatible)
Storage 2TB SSD + shared NAS
Monitors Two 27-inch

The RAM and GPU rows are the key ones. The moment when the engine editor, a local LLM helper, and image generation are all running at once comes around often.

A.1.2 Operating System and Base Setup

Item Value
OS Windows 11 Pro
Virtualization WSL2 (Ubuntu), as needed
Backup Daily, automated

WSL2 is not kept running at all times; I pull it in only when a Linux-only tool needs to run. The single most important line in this table is that backups run automatically every day.


A.2 Tool Inventory

I group the tools into five categories: engine and tools, design, art, collaboration and operations, and AI/LLM. No one person uses all five, but a game designer moves between three of them every day: design, collaboration and operations, and AI/LLM. Note that each category takes the shape of "one or two essentials plus supporting tools."

A.2.1 Game Engine and Tools

Tool Purpose
Unreal Engine 5.7 or later Main engine
Visual Studio Code
Rider C# IDE (secondary)
Perforce or SVN Code and asset version control

The engine and version control come as a pair. Even game designers must be able to handle a version control client, because the data sheets and specs all live in the same repository.

A.2.2 Design

Tool Purpose
Excel Data sheets + VBA macros
Markdown editor Specs and meeting notes
Figma UI and wireframes
Mermaid Diagrams

This is the game designer's everyday workbench. Excel is the home base for data, Markdown the home base for writing — and they are also the two points where AI tools attach most deeply. Mermaid earns its own row because, as the main text emphasizes, diagrams are the language of agreement.

A.2.3 Art

Tool Purpose
Maya / Blender 3D
Substance 3D Textures and materials
Photoshop 2D and illustration
Stable Diffusion (SDXL) / ComfyUI Self-hosted production of concepts and textures (LoRA, ControlNet)
Midjourney Early mood boards (secondary)

These are not tools a game designer uses directly, but knowing which tools the art team has on their side changes the resolution of your requests when trading concepts back and forth. Production work runs primarily on self-hosted Stable Diffusion (SDXL) / ComfyUI — assets never go outside, which protects the IP, and character LoRAs plus ControlNet let us control the same character's consistency across repeated generations. Closed tools like Midjourney serve only as a supplement for early mood boards, when we are first feeling out the project's tone; they are not used for production work where consistency and repeat control are at stake.

A.2.4 Collaboration and Operations

Tool Purpose
Collaboration tool (ClickUp) Tasks
Team messenger Real-time communication
Self-built wiki Wiki and long-lived documents
In-house web portal Unified interface (20.3)

The time axis of communication is what divides the tools. Real-time communication that needs immediacy goes to the team messenger; tasks go to the collaboration tool (our team uses ClickUp); knowledge meant to last goes to the self-built wiki. Swap the tracker for JIRA or Redmine, the wiki for Confluence or Notion, use whatever messenger you have — the flow of this book stays the same. The in-house web portal is the unified gateway that connects these three with the AI tools on one screen, and 20.3 covers it in detail.

A.2.5 AI and LLM

Tool Purpose
Claude (Opus + Sonnet) Main LLM
GPT-4 Alternative
Whisper (self-hosted) Speech recognition (STT)
Stable Diffusion Image generation (self-hosted)
MCP servers Tool integration (20.4)

Claude is the main model, with GPT-4 kept for cross-checking and as an alternative. The Whisper and Stable Diffusion rows carry a principle: sensitive audio and images are processed self-hosted, never sent outside. The MCP servers are the glue that fits these tools into the workflow; 20.4 explains the structure.


A.3 atom Inventory

An atom is a fragment of knowledge — the "minimum unit of decision" covered in the main text, dropped into a file. The table below shows how those atoms are distributed across domains, a snapshot as of May 2026. Look at where the decisions cluster rather than at the absolute counts. Where decisions cluster is where the project is thinking hardest.

Category atom Count Notes
combat 47 Combat system decisions
narrative 38 Five narrative layers
ui 31 UI and HUD
balance 28 Balance values
level 22 Level design
character 19 Characters and voice_profile
meta·governance 18 Procedures and rules
qa·integrity 16 Verification
content 14 Content production
operations 14 Operations workflows
external_reference 12 External references
economy 11 Economy and resources
Other 34 Classification in progress

The distribution — combat thickest, narrative right behind — directly reflects the character of this project: a combat-centered MMORPG that nonetheless refuses to give up its narrative weight. The "Other: 34" row holds new decisions whose categories are not yet settled; when that cell grows too large, it is a signal that the classification scheme is due for an overhaul. The total as of May 2026 is 304.


A.4 Verification and Operations Materials

Even with tools and knowledge in place, quality slides downhill without a mechanism that confirms they are actually working. This section shows that mechanism in two forms: reports produced on a regular cadence, and decision cards that leave decisions traceable after the fact.

A.4.1 Reports

Report Frequency
Daily build report Daily
Alpha gap report Weekly (10.3)
Sprint quality report Biweekly
Milestone QA report Every milestone
Quarterly retrospective Quarterly

The frequency is the report's character. What runs daily is a state check; weekly and biweekly are trend checks; milestone and quarterly are direction checks. Where AI contributes most is drafting the reports that repeat — the daily and weekly ones — and 10.3 covers that case.

A.4.2 Decision Cards

Quarter Decisions
Q4 2025 132
Q1 2026 156
Q2 2026 (in progress) 89
Cumulative 547

The bare fact that around 100 decisions per quarter end up as cards shows the operating principle: decisions are handled as records, not memories. The Q2 figure of 89 is a running count at mid-quarter, so it is a value in progress; by quarter's end it reaches the level of the previous quarter. The flow in which these cards accumulate and get promoted into the atoms of A.3 is this system's learning axis.


A.5 Meeting Distribution (for Reference)

Meetings are where the most time leaks away — and where the effect of AI tools is felt fastest. Below is the average distribution of meetings per quarter, grouped by category, offered as a reference for gauging the input volume of the meeting-notes system covered in 17.3. The numbers swing from quarter to quarter, so I give them as ranges.

Category Quarterly Average (Meetings)
Daily standups (daily) 65–70
Combat (battle) 35–45
Art (art) 25–30
Issues (issue) 8–15
Reviews (review) 6–10
Other (1:1 and external) 40–50

Daily standups are the most frequent, with combat-related meetings right behind. It is telling that this is the same shape as the atom distribution in A.3: meetings cluster in the same domains where decisions cluster. The more meeting-heavy the environment, the greater the payoff from automated meeting-notes processing, and 17.3 explains the concrete operation.


A.6 Notes for the Reader

Every table above is a single photograph of my environment. It is not a list to copy verbatim; please use it as a template for laying out your own environment in the same frame. With a different team size, genre, or platform, the tools, the atom distribution, and the meeting mix will all differ. What matters is not matching the items but whether the four layers — foundation → tools → knowledge → verification — run unbroken in your environment too. If one of those four layers has an empty cell, that cell is the next place to work on.

Appendix B. Tool Adoption Procedure (Generalizing from Company to Personal Use)

This appendix documents the procedure I used to bring the tools and skills I had built and operated on the company's Project A over to my personal PC and to general-purpose work. The core question is a single one: "How do I legitimately take only the skeleton of the tools I learned to build there, without infringing on the company's knowledge assets?" This appendix shows where I drew that line, what I brought over and what I left behind, and how I kept a record of those decisions.

Here is how to use this appendix. First, read the five principles in B.1 against your own situation, then walk through the procedure in B.3 once, exactly as written. After that, copy the record template in B.4 and fill it in for the tool you want to bring over. Because this involves company assets, "leaving a trail" takes priority over "moving fast," and this entire appendix is structured from that perspective.


B.1 The Five Principles of Adoption

These are the five principles I settled on before bringing any tool over. They are not a sequence but conditions that must hold simultaneously; if even one collapses, the adoption itself goes on hold. The first three are technical boundaries about "what to take," and the last two are procedural boundaries about "how to take it with a clear conscience."

Principle Description
1. No company IP included Remove company names, real names, and proper nouns
2. Take only the tool skeleton Block company domain data
3. Reconstruct for general use Rebuild around general use cases
4. Clear citation and attribution State explicitly that the tool was adopted from the company
5. Legal and HR agreement Go through the company's consent process

The line that wavers most often is number 2. You may take the algorithm and the structure (the skeleton), but the moment the company data formats that the skeleton assumed come along with it, you have taken IP. Separating skeleton from data is the real substance of adoption.


B.2 The Six Tools and Skills I Adopted

Following these principles, the tools I actually brought to my personal PC number six (as of May 2026). They share one trait — all of them handle data — and that is no coincidence. Data-processing tools are comparatively easy to split into skeleton (parsing, transformation, and visualization logic) and domain (the specific formats of company sheets).

Tool Company original Personal generalized edition
excel-reader xlsm sheet and VBA extraction General-purpose Excel processing
relation-map-gen FK relationship HTML General-purpose data relationship maps
schema-doc Markdown schema generation from sheets General-purpose schema documentation
table-creator Mass production of data tables General-purpose table generation
gdd-gen Automatic GDD (game design document) generation General-purpose document generation
gdd-export Markdown to multi-sheet xlsx conversion General-purpose xlsx conversion

Compare the middle and right columns and what generalization means becomes visible. The left side carries domain-laden names like "company sheets" and "GDD"; the right side carries names with the domain stripped out, like "general-purpose Excel" and "general-purpose documents." The company disappearing from the name is the first sign of generalization.


B.3 The Adoption Procedure

Translating the principles (B.1) into actual hand movements yields the six steps below. The two critical junctures are steps 2 and 4. If you fail to cleanly split skeleton from domain in step 2, every step after it gets contaminated; and if you skip the company's consent in step 4, the tool is unusable no matter how well you build it.

flowchart TD
    A[1. Identify the company tool] --> B[2. Separate company-dependent areas]
    B --> B1[Dependencies on company data and proper nouns]
    B --> B2[Tool skeleton: algorithms and structure]
    B1 --> C[3. Remove company dependencies + parameterize for general use]
    B2 --> C
    C --> D[4. Company consent: legal and manager]
    D --> E[5. Apply and verify on the personal PC]
    E --> F[6. Cite the source + record the adoption]
    classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545;
    classDef human fill:#fde68a,stroke:#b45309,color:#000;
    classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b;
    class B2 code;
    class D human;
    class B1,F data;

Of the six steps, the one that takes the longest is not the code work (steps 2 and 3) but step 4 — reaching agreement with the company and clearing legal review. The biggest gate is trust, not technology, which is why adoption always runs in this order: secure the agreement first, polish the code later.


B.4 The Adoption Record

An adopted tool must always come with a record. A moment may arrive later when someone asks, "Where did this tool come from, what was removed, and whose consent did it have?" Below is the record template using excel-reader as the example; copy this frame as is and fill it in for your own tool.

---
tool: excel-reader (personal generalized edition)
original_source: Project A (company)
adopted: 2026-05
permission: Company manager + legal review cleared
modifications:
  - Removed dependency on company sheet formats
  - Removed company domain functions (xlsm VBA)
  - Generalized to standard csv/xlsx processing
  - Removed all references to company names and real names
usage_in_book: Cited as tool examples in this book (Parts 1, 5, 6, 8, etc.)
---

For the date field (adopted), write a confirmed year-month like 2026-05. A loose entry like "around May 2026" reads like a blank to be filled in later, so pin down the moment you confirmed the adoption right then and there.

The two most valuable lines in this record are permission and modifications. The former proves the adoption was legitimate; the latter proves what was stripped out. With these two lines in place, a traceable basis remains even if questions are raised later.


B.5 Tools I Did Not Adopt

What I left behind matters as much as what I brought over. Here are the company tools I deliberately chose not to adopt, and why. What they have in common is that they are either core company IP or so deeply bound to the company's organizational structure that skeleton and domain cannot be pulled apart.

Tool Reason not adopted
Company combat system tools Core company IP, company-exclusive
Company narrative documentation tools Depends on the company's world setting
Company combat task force (TF) tools Depends on the company's organizational structure
Company HR and finance tools Does not fit external environments

This contrasts precisely with the adopted tools in B.2, which were all "data processing." What I brought were tools separable from their domain; what I left were tools fused with their domain. Separability determines adoptability.


B.6 For the Reader — A Pre-Adoption Self-Checklist

Finally, here are the five items you must pass yourself through before bringing a tool over. This table is a pass/fail checklist: adopt only when all five items pass, and hold off if even one fails. There is no "mostly fine." Partial passes do not work when company assets are involved.

Check item Pass criterion
Did you obtain the company's consent Explicit agreement from manager and legal
Did you clear legal review Written or recorded confirmation
Did you completely remove company IP Zero hits on the grep watchlist
Did you verify generality Confirmed working in other environments
Do you have an incident response procedure Tracing and recall paths defined

Read these five items not as five boxes to tick but as five locks. Bringing what you learned at a company into your personal assets legitimately is entirely possible — but that legitimacy holds only when all five locks are engaged.

Appendix C. Permissions and Settings Reference

This appendix is a reference that gathers in one place the permissions and settings for the tools and systems cited in the main text. The main text explains why I run things this way, but when you actually apply it to your own environment, what you need is "so which value, concretely, goes where?" This appendix fills that gap.

The reason behind each value matters more than the value itself. Rather than copying the numbers in the tables verbatim, read the short note under each item and adjust to your own team size and risk level. If you work alone, there is no need to split permission tiers; if you have no contractors, drop the contractor rows entirely.

There are two ways to use this appendix. When setting up an environment for the first time, walk through it in order from C.1 and use it as a checklist for anything missing. During operations, when an incident occurs, open C.7 (Incident Response) first, find the matching incident row, and trace back up to the prevention items above it.


C.1 LLM API Permissions

LLM API keys tie directly to cost, and an exposed key turns into a financial incident immediately. That is why key management and permission tiers come first.

C.1.1 Key Management

Key Storage
Anthropic API environment variable + 1Password
OpenAI API environment variable + 1Password
Self-hosted internal to the company

Inject keys through environment variables, not in code, and keep the originals in a secrets manager (e.g., 1Password). The most common incident is pushing a key to git embedded in code, so putting keys in git is prohibited without exception.

C.1.2 Permission Tiers

User Permission
Directors and seniors full (responsible for running the cost cap)
Regular members per-task cap
Contractors one-time, per task only

Permissions are divided not by trust but by the size of the responsibility. Whoever holds full permission also carries the responsibility of managing the cost cap. Contractors get access opened once per task, revoked when the task ends.


C.2 Tool Settings Standards

Recommended settings differ by tool, but the core is separating analysis work from creative work. Analysis must be reproducible; creative work needs variety.

C.2.1 Claude Code

Below is an example base configuration for Claude Code. Line by line: pin the model, turn on extended thinking, set a token ceiling, enable auto-update, and hook a memory-injection script to prompt submission.

{
  "model": "claude-opus-4-8",
  "extended_thinking": true,
  "max_tokens": 100000,
  "auto_update": true,
  "hooks": {
    "UserPromptSubmit": ["~/.claude/hooks/inject_memory.py"]
  }
}

The model value is only an example. Model names change with each generation (this example is as of the time of writing), so do not copy it verbatim — check the latest available name with /model and use that. Even when the names change, the workflow skeleton in this book keeps working (see Appendix K).

The script hooked to hooks.UserPromptSubmit automatically inserts the relevant memory fragments every time a prompt is submitted. This memory-injection mechanism is covered in detail in Part 24.

C.2.1.1 Tool Permission Schema (allow / deny)

Inside the same settings.json, permissions live in their own permissions block. Tools the AI may run automatically without human approval go in allow; tools where a single incident would be fatal, so automatic execution must be blocked, go in deny. The notation is Tool(command pattern), and :* means "every call that starts with that command."

{
  "permissions": {
    "allow": [
      "Bash(ls:*)",
      "Bash(git status:*)",
      "Bash(git diff:*)",
      "Read(*)",
      "Grep(*)"
    ],
    "deny": [
      "Bash(rm -rf:*)",
      "Bash(git push --force:*)"
    ]
  }
}

Commands with nothing to undo — reads and searches (Read, Grep) and status queries (git status, git diff) — are auto-approved via allow to cut down approval-popup fatigue. Conversely, commands where one mistake is unrecoverable, like rm -rf and git push --force, go into deny no matter how wide you make the auto-approval range.

There are four operating principles.

Principle Details
Start with a whitelist Start auto-approval at the minimum and add to allow only when needed
Explicitly block dangerous commands rm -rf and git push --force go in deny, no exceptions
Regular cleanup Re-review allow every quarter and trim unused permissions
Separate by domain Split global permissions from project permissions so the home PC and the work PC carry different policies

The allow list is not a static setting but a trace of accumulated work. It grows as your repeated tasks grow, so it pays to pair it with a quarterly cleanup cycle. The background of this permission practice is covered in detail in Part 1, Chapter 3.

C.2.2 In-House Tools

Setting Recommended value
LLM temperature (analysis) 0
LLM temperature (creative) 0.7
Cache TTL 1 hour
Cost cap (daily) defined per tool
Backup cycle daily

Analysis calls run at temperature 0 so that the same input produces the same output. Verification, lint, and classification — work whose results must not wobble — belong here. Idea divergence and draft generation, by contrast, get a variety of around 0.7. Do not set a single standard for the cost cap; define it per tool, because call frequency and token consumption differ from tool to tool.


C.3 Git Permissions

Branch Permission
main only directors and seniors push
feature/* all members
protected branches mandatory code review

The main branch blocks direct pushes; every change comes in from a feature branch through code review. Force-push overwrites collaboration history, so it is banned; an exception is made only when unavoidable incident recovery requires it, and only by agreement between the director and the code lead.


C.4 File System Permissions

Document folders divide permissions along the Layer structure (L0–L4). The higher up you go, the wider the blast radius, so write permission narrows; the lower you go, the more distributed the work, so write permission widens.

Folder Permission
docs/L0_vision/ director write, everyone read
docs/L1_systems/ domain director write, everyone read
docs/L2_content/ owner read/write
docs/L4_meta/ everyone write
team_memory/per-user/ owner-only read/write

Vision (L0): only the director writes, everyone reads. Systems (L1): the domain director writes. Content (L2): the owner writes; meta and scratch (L4): anyone writes. Personal memory is accessible only to its owner. The Layer structure itself is covered in Part 6.


C.5 Backup and Recovery

Asset Backup
git repo git itself + remote backup
sheets (Excel) git + daily backup
user data DB backup (server standard)
meeting notes and decisions git
memory daily automatic sync

Backup paths differ by asset type, but the principle is one: the harder an asset is to recover once lost, the more redundantly you keep it. For text assets (meeting notes, decisions, code), git is the backup; binaries and server data get separate backups. Set the recovery time objective (RTO) at four hours or less, and adjust that value to the downtime your team can tolerate.


C.6 Security

Area Rule
Sensitive data to external LLMs placeholders or self-hosting
Payment and personal information never send to an LLM
Quoting external material source attribution + legal review
User data protection anonymization + GDPR compliance

The rule that is easiest to keep and broken most often is "do not send sensitive data to an external LLM." When work is urgent, the temptation to paste real data as-is is strong. Payment and personal information are no-send without exception; when analysis is needed, substitute placeholders or use a self-hosted model.


C.7 Incident Response

Incident Response
Wrong information sent out due to LLM hallucination immediate retraction + report
Cost cap exceeded automatic block + review
Copyright incident stop use within 1 hour + legal
Security incident (key exposure) rotate the key immediately + review usage history
Data loss restore from backup + incident analysis

With incidents, stopping fast often matters more than preventing. Every response in the table follows the same order: stop first, analyze second. If a key is exposed, rotate the key before asking why, then review the usage history. If cost exceeds the cap, block automatically, then review. Write these response procedures down as a document and drill them regularly, so they run without hesitation when a real incident hits.

Appendix D. R&D Document Naming and Frontmatter Standard

A generalized version of the R&D document naming and frontmatter standard (_NAMING_FRONTMATTER_STANDARD) used on Project A at my company.


D.1 Naming Standard

D.1.1 atom

<category>_<topic>_<subtopic>.md

Examples:
combat_global_cooldown_constant.md
narrative_voice_profile_K_007.md
ui_button_primary_style.md

snake_case, with a category prefix.

D.1.2 Decision Cards

D<YEAR>_Q<QUARTER>_<NUMBER>.md

Example:
D2026_Q2_017.md

Year, quarter, number.

D.1.3 Meeting Notes

<category>_<YYYY-MM-DD>[_<seq>].md

Examples:
95_BattleTF_2026-05-18.md
art_review_2026-05-18_1.md
art_review_2026-05-18_2.md

D.1.4 Specs

spec_<topic>.md

Examples:
spec_combat_global_cooldown.md
spec_guild_attendance.md

D.1.5 Reports

report_<period>_<type>.md

Examples:
report_W21_alpha_gap.md
report_Q2_user_voice.md

D.2 Frontmatter Standard

D.2.1 atom

---
name: combat_global_cooldown_constant
description: Defines the standard global cooldown value for the combat system
type: atom
category: combat
status: active
priority: P0
related_atoms:
  - combat_skill_cooldown_rule
  - combat_healing_skill_cooldown_exception
created: 2026-05-18
last_modified: 2026-05-18
related:
  derives_from: [combat_design_principle]
  affects: [combat_skill_cooldown_rule, ui_skill_cooldown_indicator]
---

D.2.2 Decision Cards

---
decision_id: D2026_Q2_017
title: Unify the combat global cooldown at 0.5 seconds
type: system_change
status: active
created: 2026-05-18
created_by: Team member A
approved_by: Minsoo Lee
scope:
  - combat_system
affected_atoms: [...]
implementation:
  target_build: 2026-05-18
verification:
  layer_1: passed
  layer_2: passed
  layer_3: pending
---

D.2.3 Meeting Notes

---
type: meeting_note
category: battle
date: 2026-05-18
attendees: [Team member A, Team member B, Minsoo Lee]
related_atoms: [...]
---

D.2.4 Specs

---
title: Guild attendance feature spec
type: spec
priority: P1
target_milestone: MS2
---

D.3 Required vs. Optional Fields

D.3.1 Required Fields

Document type Required
atom name, description, type, category, status
Decision card decision_id, title, type, status, created, scope
Meeting notes type, category, date, attendees
Spec title, type, priority

D.3.2 Optional Fields (Nice to Have)

Document type Optional
atom related, last_modified, priority
Decision card rationale, related_decisions, verification
Meeting notes related_atoms, sub_topic
Spec target_milestone, related_atoms

D.4 Automated Lint Checks

# frontmatter_lint.py

for file in glob("**/*.md"):
    fm = parse_frontmatter(file)
    if not fm:
        warn(f"{file}: missing frontmatter")

    doc_type = infer_type_from_filename(file)
    required = REQUIRED_FIELDS[doc_type]

    for field in required:
        if field not in fm:
            warn(f"{file}: missing required field {field}")

Runs automatically at build time. Violations trigger an alert.


D.5 Preventing Naming Collisions

Area Prevention
atom name globally unique
Decision ID unique within a quarter
Meeting ID date + seq
Filename unique within a folder

Naming collisions are blocked automatically.


D.6 Change Procedures

D.6.1 Renaming an atom

1. Create an atom under the new name
2. Update every wikilink to the old atom with the new name (automated)
3. Mark the old atom deprecated + redirect
4. Move it to _archive after one month

Rushed renames risk damaging your records.

D.6.2 Changing the Frontmatter Standard

1. Propose the reason for the change (decision process)
2. Migration script for all existing documents
3. Update the build lint
4. Notify the team

D.7 Notes for the Reader

This standard reflects my environment. Adjust it to fit your own. The essentials:

Essential Why
Naming consistency Search and automation
Frontmatter standard Tool-friendly
Required/optional split Writing burden ↓
Automated lint Enforces the standard
Change procedures Protects your records

Appendix E. MCP Server Catalog (A Game Design Perspective)

MCP (Model Context Protocol) is the channel through which an LLM connects to external tools and data in a standardized way. Part 20 covered project management MCP servers, but the MCP servers you can pull into a game design workflow go far beyond that. This appendix is a catalog that gathers those candidates in one place and assigns priorities for the order in which they are worth adopting.

The point of this catalog is not "install all of these" but "know where to look when you need one." If you attach several MCP servers at once, you cannot tell which one is causing problems. Follow the adoption cycle in E.4 and add them one at a time.

Here is how to use it. At first, look only at the P0 list in E.2.1. Once the basics are in place, move on to E.2.2 (P1); when your team develops specific needs, consider E.2.3 (P2) or E.3 (building your own). If cost is a concern, read E.5 first; if you want to prepare for failures, start with E.6.


E.1 The Four Application Areas of MCP

MCP servers fall into four broad areas depending on what they connect to. Most of the tools a game designer touches every day fit inside them.

Area MCP Servers
Project management ClickUp, JIRA, Linear
Documents Confluence, Notion, Google Drive
Collaboration Team messengers (Slack, Discord, etc.)
Data Excel, Google Sheets, DB

Project management leads to tasks and schedules, documents to design docs and wikis, collaboration to team communication, and data to balance and item sheets. Pin down which area the tools your team already uses belong to, and the adoption candidates narrow themselves naturally.


E.2 Recommended MCP Servers (Game Design Priorities)

I assigned the priorities by one criterion: does work stall without it? P0 is the foundation for almost all work, P1 makes things much easier to have, and P2 is a choice that depends on your team's situation.

E.2.1 P0 — Adopt First

Server Use Notes
Filesystem MCP Local file access Foundational
Git MCP Change tracking Essential
Team messenger MCP Team communication Recommended
Collaboration tool MCP (ClickUp, JIRA, etc.) Tasks Your company's tool

Filesystem and Git come first because they are the foundation that lets the LLM read materials and follow change history. The team messenger MCP pulls in team context, and for tasks, connect whatever tool your company already uses — whether that is ClickUp or JIRA.

E.2.2 P1 — Add Next

Server Use
Wiki MCP (Confluence, Notion, etc.) Wiki
Google Drive MCP Externally shared materials
Excel MCP Direct sheet queries
Mermaid MCP Diagram rendering

Once P0 is stable, expand toward documents and data. Excel MCP in particular lets the LLM query balance sheets directly, which makes it highly useful in game design. Mermaid MCP renders design diagrams on the spot, so it never breaks the documentation flow.

E.2.3 P2 — Optional

Server Use
Discord MCP User community
GitHub MCP External collaboration
Linear MCP Alternative task tool
Notion MCP Alternative wiki

P2 servers are alternatives or built for specific situations. Attach Discord if you run a user community, and GitHub if you collaborate with outside parties often. Linear and Notion are substitutes for tools you have already adopted, so there is no need to install duplicates.


E.3 Game-Specific MCP Servers (Built by the Author)

Where commercial MCP servers leave a gap, I build my own. Below are the MCP servers I developed in-house to fit my game design workflow. All of them exist to query the systems covered in the main text — atoms, decision cards, and meeting notes — directly from the LLM.

Server Use
Atom MCP atom search and lookup
Decision Card MCP Decision card lookup and creation
KPI Dashboard MCP Dashboard data
Meeting Notes MCP Meeting notes search

These four handle in-house assets that commercial tools do not cover: knowledge atoms, decision history, and meeting notes. Building your own is a heavy lift, so defer it to the last stage of the E.4 cycle and start only once it is clear that no commercial MCP server can fill the gap.


E.4 The MCP Adoption Cycle

Attach MCP servers all at once and it becomes hard to isolate the cause of a problem. The cycle below stretches one principle — one at a time, and the next only after the last is stable — along a time axis.

flowchart LR
    A["Week 1
Pilot a single Filesystem server"] --> B["Month 1
Add Git + team messenger"] B --> C["Month 3
Run 5–7 servers stably"] C --> D["Month 6
Consider building your own MCP"] classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef human fill:#fde68a,stroke:#b45309,color:#000; class A,B,C code; class D human;

There is exactly one core rule: never adopt 5 at the same time. Each time you attach a new MCP server, spend a few days watching whether that one runs stably before moving on to the next.


E.5 MCP Operating Costs

Server Cost
External MCP (open source) Infrastructure only
Self-hosted Infrastructure + operations
Commercial MCP Monthly subscription

The cost structure splits three ways. Open-source MCP servers cost only the infrastructure to run them, self-hosting adds the cost of operations staff on top, and commercial MCP servers charge a subscription. Running 8–10 servers, the monthly cost comes to roughly $50–200 by my estimate, but it varies widely with your configuration — treat it as a direction, not a figure.


E.6 Incident Response

Incident Response
MCP server outage Run fallbacks for critical servers
Permission incident (wrong data modified) Default to read-only
Data leak Self-host sensitive data
Cost explosion Cap + monitoring

MCP wires external tools directly into the LLM, so a single bad write can corrupt real data. That is why the default stays read-only, and write permission opens only on servers that truly need it. Prepare fallbacks for critical servers in case of outages, and run any MCP server that handles sensitive data on self-hosted infrastructure rather than externally. Contain cost with a cap and monitoring together.


E.7 Pre-Adoption Self-Checklist

The preceding sections covered what to attach, in what order, and at what cost. This table collects the items you should pass yourself through right before actually attaching a single MCP server. Instead of rereading the catalog from the top, recheck just these five lines every time you add a new MCP server. Each of the five items compresses a core rule from an earlier section into one line.

Check Item Pass Criterion Source Section
Which area does it belong to Clearly falls under one of project management, documents, collaboration, or data E.1
Is it the right priority for now P1 only after P0 is stable, then P2 — in that order E.2
Are you attaching one at a time No adopting several at once E.4
Are permissions minimal Read-only by default; write only on servers that truly need it E.6
Is there a cost ceiling A cap and monitoring are both in place E.5

Of the five items, the one most often skipped is "one at a time." Put several MCP servers up at once and, when a problem appears, you cannot tell which server is to blame. Attach an MCP server only when it passes all five lines; if it trips on even one, push that server to the next cycle.

Appendix F. Case Index (Company / Personal PC)

An index of the cases that appear in this book, split into company environment vs. personal-PC environment. A resource for quickly finding the cases closest to your own setting.


F.1 Company-Environment Cases (MMORPG Studio A, Project A)

F.1.1 System Cases

Case Where It Appears
Running CombatBalance and CombatFormula 8.1
Economy Machinations Pilot 8.2
Damage Simulator (2008–) 8.3
Procedural Level Design Master 7.1
The Behavior Tree editor 7.2
The dungeon and field pattern library 7.3
HUD Layout v3 9.1
The Skill UI 6-column decision 9.2
The NarrativeDocs 5-layer structure 5.1
voice_profile + voice_lint 5.2, 5.4
proj_city_hunting_generator 6.2
NPC Persona/Squad 6.3

F.1.2 Operations Cases

Case Where It Appears
Running 95_BattleTF 16.1
97_DevGuide collaboration 16.2
The 17.x meeting-notes system All of Part 17
Alpha Gap Report 10.3
decision_validation 3-layer 10.2
Running 304 atoms 20.1
Per-member memory 20.2
The design portal 20.3

F.1.3 Organizational Cases

Case Where It Appears
Vision and roadmap for a mid-size (10–50 person) team 19.1
Delegation as a design director 19.2
Conflict management and team culture 19.3
Running meetings (a leader's perspective) 19.4
Upward communication (PD/CEO) 19.5
AI adoption strategy 19.6
Governance (prompts, hallucination, cost, legal, ethics) All of Part 22

F.2 Personal-PC Environment Cases

Cases I experienced firsthand in my personal-PC environment (at home).

F.2.1 Borrowing Tools

Case Where It Appears
Borrowing six tools, including excel-reader Appendix B
The JIT atom injection system (personal-PC infrastructure)
Personal-PC slash commands (book-capture and others) (personal-PC infrastructure)

F.2.2 Writing This Book Itself

The process of writing this book is itself a case of AI in practice.

Area Application
Mass-producing chapter drafts LLM (Claude)
IP protection (company → anonymized) grep watchlist + rules
Source tracking Stated explicitly whenever the company environment is cited
The produce → review → polish cycle Entered review mode after the May production run

F.3 A Guide to Using the Cases

F.3.1 Readers in a Company Environment

The company-environment cases (a mid-size 10–50 person team, an MMORPG, live ops) apply to companies of similar scale and domain.

Your Environment Best-Fit Cases
Mobile MMORPG studio Almost every case
PC MMORPG Adjust the mobile cases in Part 14
Indie games Scale down the cases meant for mid-size (10–50 person) teams and up
Live-ops games Part 15 plus the operations cases

F.3.2 Readers in a Personal Environment

In a personal setting (one or two people, or a hobby project), borrow the company cases in simplified form.

Area Simplification
Meeting system Unnecessary for one person; use your own notes
TF operations Unnecessary for one person
Decision cards Only for big decisions
atoms and wikilinks Use them in earnest (valuable even solo)

F.4 IP Handling When Citing Cases

Every company case in this book is anonymized.

Original Anonymized As
Company name MMORPG studio A
Project Project A
Team members' real names Team members A, B, C
In-game proper nouns Fictionalized (Kingdom X, character K_001, etc.)
Numbers Disguised (the ratios are real)
Company tool names proj_* (e.g., proj_city_hunting_generator)

F.5 Non-Game Role Index — Finding the "Beyond Games" Boxes

A reverse index for readers who work outside games — in planning roles, as PMs, or as office professionals in general. The "Beyond Games" box at the end of each chapter is the bridge that carries that chapter's workflow into work that has nothing to do with games. If the game-domain body text feels like a lot, you are welcome to open these boxes first and enter through cases from your own line of work. This index is the anchor for the "General-Role Path" (Parts 1–2 → 17 → 16 → 18 → Parts 21–22) and the 90-minute express course (17.1 → 16.2 → 22.1 → 21.1).

F.5.1 Process and Collaboration (the Most Direct Transfer Beyond Games)

Chapter What the "Beyond Games" Box Carries Over
16.1 Isolating a flood of incoming work in a temporary workspace and absorbing only the results as canon
16.2 Sorting a one-line request into three tracks: agreement, defect, and schedule
16.3 Framing deliverables in the medium that fits each discipline and stakeholder
17.1 Making meeting notes flow into the four decision fields (what, who, why, next)
17.2 A pipeline that extracts decisions and actions from meeting notes
17.3 Classifying and syncing meeting decisions
17.4 Automating meeting summaries and follow-up tracking
18.1 Giving decisions a permanent address, an owner, and a rationale — and searching past decisions first
18.2 Classifying how far a single decision's impact propagates
18.3 A before-and-after change-tracking workflow
18.4 Checking a document change's impact scope with search

F.5.2 Self-Improvement and Governance

Chapter What the "Beyond Games" Box Carries Over
21.1 Making the retrospective the starting point of self-improvement
21.2 Promoting recurring patterns from retrospectives into rules
21.3 Closing the improvement loop
22.1 Putting context, format, hallucination blocking, and verification into a one-page work order (the prompt)
22.2 Layered defense for hallucination and safety
22.3 Managing AI cost honestly
22.4 Copyright and ethics checks

F.5.3 Leadership

Chapter What the "Beyond Games" Box Carries Over
19.1 Setting the vision and delegating
19.2 Conflict management and meeting leadership
19.3 An organization's AI adoption strategy

This index collects only the "Beyond Games" boxes that actually exist in the body text (22 of them, as of June 2026). Chapters without a box depend heavily on the game domain and resist a one-to-one transfer; rather than forcing it, I recommend entering through the "General-Role Path" chapters above.

Appendix G. Operations Script Casebook

This appendix is a casebook that gathers in one place the operations automation scripts mentioned in the main text. The main text explained why each script was needed as part of the larger flow, but when you actually sit down to build something similar, you need a map that shows at a glance which scripts exist and how they group by role. This appendix is that map.

For each script I list its name, a one-line description, and the section of the main text where it appears. For the core scripts that generalize cleanly — the format lint (G.1.1), the integrity check (G.2.1), the relation graph (G.3.1), and the cost tracker (G.7.1) — and for the test and hook examples in G.8, I included real code: written fresh as general skeletons unrelated to any company material and verified to run as is. The sample inputs, outputs, and exit codes are all values I confirmed by actually running them. The remaining entries carry only a name, a role, and the linked section in the main text; I explain the reason honestly in G.9. Use the real-code entries as models to build implementations that fit your own environment.

Here is how to use it. First decide the nature of the task you want to automate (verification, report generation, or synchronization), then open the matching section (G.1 through G.7). Pick the closest script there, follow the section number in parentheses back to the main text to check the context and design intent, and finally test your own script against the operating principles in G.8.

Grouped by role, the full set looks like this.

flowchart TD
    G1["G.1 Meeting notes and decision automation"] --> META["Meta operations
(knowledge accumulation)"] G2["G.2 Verification and lint"] --> QA["Quality gate"] G3["G.3 Impact tracking"] --> QA G4["G.4 Automated report generation"] --> REPORT["Reporting and visibility"] G5["G.5 Synchronization"] --> META G6["G.6 LLM integration"] --> AI["AI assistance"] G7["G.7 Cost and operations"] --> AI classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545; classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764; classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; class G1,G2,G3,G4,G5,G6,G7 code; class AI ai; class META,REPORT data;

G.1 Meeting Notes and Decision Automation

This group of scripts keeps decisions made in meetings from scattering, so they accumulate as knowledge assets. It runs as one continuous line from meeting-note linting through atom extraction to formal promotion.

G.1.1 meeting_lint.py

Checks whether a meeting note follows the required format (mandatory frontmatter and mandatory sections). A note with a broken format breaks the automated extraction downstream, so this blocks it at the door (17.2.2).

Below is a general skeleton, unrelated to any company material. It uses only the standard library (just sys) and runs as is. It checks that a Markdown meeting note has all the frontmatter keys (the block wrapped in ---) and all the body section headings (## ...). If anything is missing, it emits violations and exits 1; if everything is present, it exits 0.

#!/usr/bin/env python3
"""meeting_lint.py

Checks whether a Markdown meeting note follows the required format.
- Whether the frontmatter (--- block) contains all required keys.
- Whether the body contains all required section headings (## ...).
If anything is missing, print the violations and exit 1; otherwise exit 0.
Uses only the standard library.

Usage:
    python meeting_lint.py meeting.md
"""
import sys

REQUIRED_FRONTMATTER = ["type", "date", "category", "attendees"]
REQUIRED_SECTIONS = ["## Agenda", "## Decisions", "## Action Items", "## Next Meeting"]


def lint(text):
    """Take the meeting-note text and return the list of missing items (violations)."""
    violations = []

    # Frontmatter: if the first line is ---, treat everything up to the next --- as frontmatter.
    lines = text.splitlines()
    front = []
    if lines and lines[0].strip() == "---":
        for line in lines[1:]:
            if line.strip() == "---":
                break
            front.append(line)
    front_keys = [ln.split(":", 1)[0].strip() for ln in front if ":" in ln]
    for key in REQUIRED_FRONTMATTER:
        if key not in front_keys:
            violations.append({"kind": "frontmatter", "missing": key})

    # Sections: check that each heading line appears verbatim in the body.
    body_lines = [ln.strip() for ln in lines]
    for section in REQUIRED_SECTIONS:
        if section not in body_lines:
            violations.append({"kind": "section", "missing": section})

    return violations


def main(argv=None):
    argv = sys.argv[1:] if argv is None else argv
    if len(argv) != 1:
        sys.stderr.write("Usage: python meeting_lint.py meeting.md\n")
        return 2
    with open(argv[0], encoding="utf-8") as f:
        violations = lint(f.read())

    for v in violations:
        print(f"[VIOLATION] {v['kind']}: {v['missing']}")
    if violations:
        sys.stderr.write(f"[FAIL] {len(violations)} format violations\n")
        return 1
    sys.stderr.write("[PASS] format OK\n")
    return 0


if __name__ == "__main__":
    sys.exit(main())

The two constants are the check criteria. For example, feed it a meeting note whose frontmatter is missing attendees and whose body lacks ## Next Meeting, and it catches both, like this, with exit code 1.

[VIOLATION] frontmatter: attendees
[VIOLATION] section: ## Next Meeting

G.1.2 decision_parser.py

Reads the "Decisions" section of a meeting note and automatically extracts knowledge-atom candidates. It replaces the work a person used to do by hand, copying entries one by one (17.2.3).

G.1.3 promote.py

Promotes atoms in pending (awaiting review) state to the formal atom folder. It puts a human review gate between automated extraction and formal assets (17.2.6).


G.2 Verification and Lint

These are quality gates that automatically catch data and content that break the rules. The machine filters out the consistency errors that human eyes easily miss.

G.2.1 integrity_check_id_uniqueness.py

Verifies that data entry IDs are unique, with no duplicates. An ID collision is an accident that only blows up at runtime, so this stops it at the data stage (10.1.2).

Below is a general skeleton, unrelated to any company material. It uses only the standard library (csv, json, sys, argparse), and you can save it and run it right away. The input is a simple format any game data set would plausibly have: a CSV with an id column.

#!/usr/bin/env python3
"""integrity_check_id_uniqueness.py

Checks that the id column of a CSV is unique.
- If duplicate ids exist, print the violation list and exit 1.
- If all ids are unique, exit 0.
Uses only the standard library.

Usage:
    python integrity_check_id_uniqueness.py data.csv
    python integrity_check_id_uniqueness.py data.csv --id-column quest_id
"""
import argparse
import csv
import json
import sys


def find_duplicate_ids(rows, id_column):
    """Find duplicate id_column values in rows (a list of dicts).

    Returns: a list of violations. Each entry has the form
    {"id": value, "row_numbers": [1-based row number, ...]}.
    The header counts as row 1, so the first data row is 2.
    """
    seen = {}  # id value -> list of row numbers where it appeared
    for index, row in enumerate(rows):
        row_number = index + 2  # data rows start after the header (row 1)
        key = row.get(id_column, "")
        seen.setdefault(key, []).append(row_number)

    violations = []
    for key, row_numbers in seen.items():
        if len(row_numbers) > 1:
            violations.append({"id": key, "row_numbers": row_numbers})
    violations.sort(key=lambda v: v["row_numbers"][0])
    return violations


def load_rows(csv_path):
    with open(csv_path, newline="", encoding="utf-8") as f:
        return list(csv.DictReader(f))


def main(argv=None):
    parser = argparse.ArgumentParser(description="Check CSV id uniqueness")
    parser.add_argument("csv_path", help="path to the CSV file to check")
    parser.add_argument("--id-column", default="id", help="column name to use as the id (default: id)")
    args = parser.parse_args(argv)

    rows = load_rows(args.csv_path)
    violations = find_duplicate_ids(rows, args.id_column)

    # G.8 output standard: emit violation_list as JSON on stdout.
    print(json.dumps({"violation_list": violations}, ensure_ascii=False, indent=2))

    if violations:
        sys.stderr.write(f"[FAIL] {len(violations)} duplicate ids found\n")
        return 1
    sys.stderr.write("[PASS] no duplicate ids\n")
    return 0


if __name__ == "__main__":
    sys.exit(main())

Sample input (data.csv):

id,name
Q001,First Commission
Q002,The Lost Norigae
Q001,First Commission (duplicate)

Running it produces the following. Q001 appears twice, on rows 2 and 4, so one violation is caught and the exit code is 1.

{
  "violation_list": [
    {
      "id": "Q001",
      "row_numbers": [2, 4]
    }
  ]
}

G.2.2 voice_lint.py

Checks NPC dialogue for voice consistency (speech style and personality). It catches the mismatch where the same character speaks differently from one chapter to the next (5.2, 5.4).

G.2.3 visual_regression.py

A regression-check script that compares assets (art, UI, and so on) after a change to see whether unintended visual changes crept in (12.1.5).


G.3 Impact Tracking

This group of scripts tracks what shakes loose when you change one thing. It follows the links among documents, decisions, and assets to show the blast radius of a change.

G.3.1 wikilink_graph.py

Scrapes the wikilinks between documents ([[target]]) and automatically builds a link graph. It lets you see at a glance which document references which (24.3.4).

Below is a general skeleton, unrelated to any company material. It uses only the standard library (os, re, json, argparse). It reads the .md files in one folder, treats each file name (minus the extension) as a node and each [[...]] link as an edge, and outputs both an adjacency list and Mermaid diagram code.

#!/usr/bin/env python3
"""wikilink_graph.py

Builds a graph of the [[Wikilink]] connections among .md documents in a folder.
- Nodes: file names without the extension.
- Edges: [[target]] notations in the body. For [[target|label]], only the target counts.
Uses only the standard library.

Usage:
    python wikilink_graph.py ./docs
    python wikilink_graph.py ./docs --format mermaid
"""
import argparse
import json
import os
import re
import sys

WIKILINK = re.compile(r"\[\[([^\]|#]+)")  # [[target]] / [[target|label]] / [[target#anchor]]


def extract_links(text):
    """Extract link target names from the body, in order of appearance, deduplicated."""
    result = []
    for match in WIKILINK.findall(text):
        target = match.strip()
        if target and target not in result:
            result.append(target)
    return result


def build_graph(doc_dir):
    """Scan the folder's .md files into a {doc name: [link target, ...]} adjacency list."""
    graph = {}
    for name in sorted(os.listdir(doc_dir)):
        if not name.endswith(".md"):
            continue
        node = name[:-3]
        path = os.path.join(doc_dir, name)
        with open(path, encoding="utf-8") as f:
            graph[node] = extract_links(f.read())
    return graph


def to_mermaid(graph):
    """Convert the adjacency list into a Mermaid flowchart code string."""
    lines = ["flowchart LR"]
    for node, targets in graph.items():
        if not targets:
            lines.append(f'    {_id(node)}["{node}"]')
        for target in targets:
            lines.append(f'    {_id(node)}["{node}"] --> {_id(target)}["{target}"]')
    return "\n".join(lines)


_ID_CACHE = {}


def _id(name):
    """Mermaid node ids must be ASCII. Korean names get short ASCII ids n1, n2, ...
    in first-seen order, and the original name is preserved in the [...] label."""
    if name not in _ID_CACHE:
        _ID_CACHE[name] = "n%d" % (len(_ID_CACHE) + 1)
    return _ID_CACHE[name]


def main(argv=None):
    parser = argparse.ArgumentParser(description="Wikilink graph builder")
    parser.add_argument("doc_dir", help="folder containing the .md documents")
    parser.add_argument("--format", choices=["json", "mermaid"], default="json")
    args = parser.parse_args(argv)

    graph = build_graph(args.doc_dir)
    if args.format == "mermaid":
        print(to_mermaid(graph))
    else:
        print(json.dumps(graph, ensure_ascii=False, indent=2))
    return 0


if __name__ == "__main__":
    sys.exit(main())

Sample input (three files in a docs/ folder). The Korean file names — 세계관 (worldbuilding), 지역_한양 (region: Hanyang), 세력_의금부 (faction: Uigeumbu) — are kept as is, since handling non-ASCII names is exactly what this example demonstrates:

docs/세계관.md     body links to [[지역_한양]] and [[세력_의금부]]
docs/지역_한양.md  body links to [[세력_의금부]]
docs/세력_의금부.md  no links

Run it with --format mermaid and you get the diagram code below. Nodes are processed in file-name order (세계관 → 세력_의금부 → 지역_한양), and the original Korean names survive inside the labels. You can see at a glance which document reaches where, and which one is the endpoint (세력_의금부).

flowchart LR
    n1["세계관"] --> n2["지역_한양"]
    n1["세계관"] --> n3["세력_의금부"]
    n3["세력_의금부"]
    n2["지역_한양"] --> n3["세력_의금부"]

G.3.2 decision_impact.sh

Analyzes which documents and assets a given decision card affects. Check the blast radius before you reverse a decision (18.4.3).

G.3.3 find_skills_using.py

Finds, in reverse, the skills that use a given asset. Know what depends on an asset before you modify or delete it (11.2.4).


G.4 Automated Report Generation

These scripts bundle scattered data into reports and diagrams a person can read. Automating recurring periodic reports cuts down on manual chores.

G.4.1 alpha_gap_report_generator.py

Tallies the gap against alpha-stage targets and automatically generates a weekly report (10.3.3).

G.4.2 decision_graph_to_mermaid.py

Converts the link structure among decision cards into Mermaid diagram code. See the flow of decisions as a picture (24.2.3).

G.4.3 weekly_kpi_summary.py

Summarizes key performance indicators (KPIs) on a weekly basis (13.2).


G.5 Synchronization

These scripts keep material scattered across multiple locations efficiently aligned. Instead of copying everything every time, they pick out and sync only what changed.

G.5.1 incremental_sync.py

Syncs only the changed meeting notes rather than all of them. Full copies get slower as material accumulates, so it goes incremental (17.5.4).

G.5.2 Change Detection with git diff

An approach that uses git's diff to detect efficiently what changed. No separate tracking machinery — git itself serves as the change detector (17.5.4.1).


G.6 LLM Integration

These scripts delegate work that requires judgment — classification, invocation, and the like — to an LLM. Tasks that do not reduce cleanly to rules get handled with LLM assistance.

G.6.1 faq_classifier.py

Automatically classifies incoming FAQs by category (13.1.3).

G.6.2 meeting_classifier.py

Automatically classifies meetings into categories by their nature. It is used to fill in the category field in the meeting-note frontmatter (17.3.6).

G.6.3 prompt_library_loader.py

Loads the prompt you need from a pre-organized prompt library, so the same prompt never gets rewritten from scratch (22.1.2).


G.7 Cost and Operations

These scripts keep the automation itself from creating blind spots in cost and source tracking.

G.7.1 llm_cost_tracker.py

Tracks LLM call costs and applies a cap. It stops a cost blowup before it happens, not after (22.3.5).

Below is a general skeleton, unrelated to any company material. It uses only the standard library (json, os, argparse). It records token counts per call, computes the running cost, and emits a rejection signal (exit 2) when the cap is exceeded. The unit prices are constants in the code; swap in the actual price table of whatever model you use (the values below are placeholders for illustration).

#!/usr/bin/env python3
"""llm_cost_tracker.py

Keeps a running record of LLM call tokens and checks a daily cost cap.
- record: adds one call (input/output tokens) to the ledger file.
- If the cumulative cost exceeds the cap, exit 2 blocks the call (pre-emptive cutoff).
Uses only the standard library.

Usage:
    python llm_cost_tracker.py --ledger ledger.json --in 1200 --out 800
    python llm_cost_tracker.py --ledger ledger.json --in 1200 --out 800 --cap-usd 5.0
"""
import argparse
import json
import os
import sys

# Unit prices: USD per 1,000 tokens. Placeholder values for illustration — replace with your model's actual price table.
PRICE_PER_1K_INPUT = 0.003
PRICE_PER_1K_OUTPUT = 0.015


def cost_of(in_tokens, out_tokens):
    """Compute the cost (USD) of one call from its input/output tokens."""
    return (in_tokens / 1000) * PRICE_PER_1K_INPUT + (out_tokens / 1000) * PRICE_PER_1K_OUTPUT


def load_ledger(path):
    if os.path.exists(path):
        with open(path, encoding="utf-8") as f:
            return json.load(f)
    return {"calls": 0, "in_tokens": 0, "out_tokens": 0, "total_usd": 0.0}


def save_ledger(path, ledger):
    with open(path, "w", encoding="utf-8") as f:
        json.dump(ledger, f, ensure_ascii=False, indent=2)


def main(argv=None):
    parser = argparse.ArgumentParser(description="LLM cost tracking and cap")
    parser.add_argument("--ledger", required=True, help="path to the cumulative ledger JSON file")
    parser.add_argument("--in", dest="in_tokens", type=int, required=True, help="input tokens for this call")
    parser.add_argument("--out", dest="out_tokens", type=int, required=True, help="output tokens for this call")
    parser.add_argument("--cap-usd", type=float, default=None, help="cumulative cost cap (USD); block when exceeded")
    args = parser.parse_args(argv)

    ledger = load_ledger(args.ledger)
    this_cost = cost_of(args.in_tokens, args.out_tokens)

    ledger["calls"] += 1
    ledger["in_tokens"] += args.in_tokens
    ledger["out_tokens"] += args.out_tokens
    ledger["total_usd"] = round(ledger["total_usd"] + this_cost, 6)
    save_ledger(args.ledger, ledger)

    print(json.dumps({"this_call_usd": round(this_cost, 6), "ledger": ledger}, ensure_ascii=False, indent=2))

    if args.cap_usd is not None and ledger["total_usd"] > args.cap_usd:
        sys.stderr.write(f"[CAP] cumulative {ledger['total_usd']} USD > cap {args.cap_usd} USD — blocked\n")
        return 2
    return 0


if __name__ == "__main__":
    sys.exit(main())

Sample input and result. Starting from an empty ledger, recording 1,200 input and 800 output tokens makes this call cost 1200/1000*0.003 + 800/1000*0.015 = 0.0036 + 0.012 = 0.0156 USD.

{
  "this_call_usd": 0.0156,
  "ledger": {
    "calls": 1,
    "in_tokens": 1200,
    "out_tokens": 800,
    "total_usd": 0.0156
  }
}

Pass --cap-usd 0.01 along with it, and the cumulative 0.0156 exceeds the 0.01 cap, so exit code 2 blocks the next call. This is what "stop it before, not after" actually does.

G.7.2 source_tracker.py

Automatically records the sources of quoted and referenced material, leaving a trail you can retrace later (24.5.4).


G.8 Script Operating Principles

Making your scripts run reliably matters more than making many of them. The five principles below apply to every script above.

Principle Description
Simplicity Avoid complex libraries
Testing Unit-test every script
Output standard Standards such as violation_list (10.1.7)
Version control git
Human review gate Automation still gets human review

The last principle matters most. Automation does not replace people; it shrinks the steps that come before human judgment. Whether the script verifies, extracts, or generates, always put a gate where a person looks once before final application.

G.8.1 Unit Test Example

Rather than leave the "testing" principle as words, here is an actual test that verifies G.2.1's core function find_duplicate_ids with the standard library's unittest. There are no external dependencies, so save it as is and run python -m unittest test_integrity_check -v. The key point: a function is this easy to test only when it is separated from file I/O — which is why G.2.1 splits the check logic apart from load_rows.

# test_integrity_check.py
import unittest

from integrity_check_id_uniqueness import find_duplicate_ids


class TestFindDuplicateIds(unittest.TestCase):
    def test_no_duplicates_returns_empty(self):
        rows = [{"id": "Q001"}, {"id": "Q002"}]
        self.assertEqual(find_duplicate_ids(rows, "id"), [])

    def test_one_duplicate_reports_row_numbers(self):
        rows = [{"id": "Q001"}, {"id": "Q002"}, {"id": "Q001"}]
        self.assertEqual(
            find_duplicate_ids(rows, "id"),
            [{"id": "Q001", "row_numbers": [2, 4]}],
        )

    def test_missing_column_treated_as_empty_string(self):
        rows = [{"name": "a"}, {"name": "b"}]
        result = find_duplicate_ids(rows, "id")
        self.assertEqual(result, [{"id": "", "row_numbers": [2, 3]}])


if __name__ == "__main__":
    unittest.main()

Run it and all three tests pass.

test_missing_column_treated_as_empty_string ... ok
test_no_duplicates_returns_empty ... ok
test_one_duplicate_reports_row_numbers ... ok

----------------------------------------------------------------------
Ran 3 tests in 0.000s

OK

G.8.2 Silent Failure in Hooks (exit 0)

Of the principles above, the one that slips through most easily is hook failure handling. A hook that runs automatically before a commit or on save should be a side branch of the main task (the commit or the save). But if the hook exits nonzero because of an internal error, the main task it is attached to gets blocked outright — the auxiliary device takes the main body hostage. So an auxiliary hook, no matter what happens inside it, should only leave a warning on standard error (stderr) and return exit code 0, so the main task is never blocked. Here is the minimal form; even when an exception is raised inside, the exit code is 0.

import sys

def run_hook():
    raise RuntimeError("internal error")

def main():
    try:
        run_hook()
    except Exception as exc:
        sys.stderr.write(f"[hook] warning: {exc} — not blocking the main task\n")
    return 0  # An auxiliary hook never blocks the main task, no matter what

if __name__ == "__main__":
    sys.exit(main())

Run it and you see the warning, but the exit code is 0. A person can tell what went wrong, and the workflow is not interrupted.

[hook] warning: internal error — not blocking the main task
(exit code 0)

Use this "silent failure" only for auxiliary hooks. Verification whose whole purpose is the pass/fail verdict — like the quality gates in G.2 — must do the opposite: exit nonzero on failure (the exit 1 we saw earlier) and stop the pipeline. Even in the same hook slot, the exit-code policy flips completely depending on whether the hook is "auxiliary" or a "gate."

G.8.3 How to Notice and Recover from a Silent Failure

The exit 0 policy of the previous section comes with one price. If an auxiliary hook never blocks the main task no matter what, then flip that around: the main task keeps running fine even when the hook dies silently. A hook that runs on a side branch, like automatic context injection, can stay down for days without a single red light in your workflow. So an auxiliary hook must always carry a companion device: alongside "failure does not block," you need "a person eventually sees the failure." Drop the companion, and one day in a retrospective you notice "this atom has not come up at all lately" — and only then learn the hook has been dead for a week.

That companion is a log. Do not let the warning from the minimal form above (sys.stderr.write(...)) evaporate — drop it into a file: one line per normal call, one line with the reason per failed call. In my environment this trail accumulates in ~/.claude/hooks/_injection_log.txt (the same log the trigger verification in §21.3.4 reads). The operating loop is not grand. One lap of a three-step check-and-recover procedure is enough.

Step What you look at What you do
Detect Whether the log's recent normal-injection lines have stopped, or failure lines with the same reason keep repeating Skim the log tail once during the weekly retrospective (one auto-captured line is enough)
Isolate Whether the failure reason is a bug in the hook itself or in the input data (a broken manifest, a missing atom file) Split the two by the stderr reason string — fix the code if it is a code problem, the manifest if it is a data problem
Recover Whether the trigger produces a normal injection again After the fix, enter the intended trigger once in a new session and confirm a normal line lands in the log again (same as the trigger verification in §21.3.4)

The point is that "detect" rests not on human attentiveness but on one log file and one retrospective line. What exit 0 prevented was the interruption of the main task — not the concealment of failure. Failures surface through stderr into the log, the retrospective looks at that log on a schedule, and recovery reuses the trigger verification you already use. Only when "don't block + surface it + look regularly + revive the same way" travel as one bundle does a silent failure not harden into silent neglect.


G.9 A Note for Readers

The code in this casebook comes in two kinds. The first — G.1.1, G.2.1, G.3.1, G.7.1, and G.8 — is code I wrote fresh as general skeletons unrelated to any company material and verified to run as is. It uses only the standard library, and every sample input, output, and exit code above is a result I confirmed by actually running it. Copy-paste it and use it right away, changing only the placeholder values — the price table, the column names — to fit your environment.

The second kind, like the rest of the entries, lists only a name, a role, and the linked section in the main text. Honestly, there are two reasons I did not print these in full. First, the original company operations scripts are company IP and cannot be carried over as is. Second, much of their logic is bound to company-specific data schemas, folder structures, and decision-card formats; strip those premises away and no code remains that a general reader could use directly. So I promoted only the four that generalize cleanly — the format lint, the integrity check, the relation graph, and the cost tracker — to real code, and left the rest as skeletons. Use these four as models and build an implementation for your own environment the same way: separate the check logic from I/O, emit the violation list on standard output, and attach unit tests.

For the procedure of taking an existing tool and adapting it, see Appendix B.

Appendix H. Reusing Past Work Materials

A game designer who has worked long enough accumulates decades of work materials. Meeting notes, decision records, retrospectives, study notes, lessons pulled from failures. This appendix covers how to put those materials back to work on a new project. The core tension is a single one: much of the material is company IP and cannot be moved freely, yet mixed into it is personal learning that applies anywhere. Drawing the line between the two is where reuse begins.

How you use this appendix depends on where you stand. If you are about to pull old materials into a new project, follow H.2 (the separation principle) and H.3 (the procedure) in order. If you worry about causing an incident while moving things over, read H.5 (the five pitfalls) first and steer clear of them. If you are early in your career and have little material stacked up yet, see H.6 and decide what to keep — and how — starting now.

The principles here are not some grand theory of asset management. They compress into one sentence: keep the concrete at the company, take only the abstract patterns. Everything else is how to apply that sentence to real situations.


H.1 The Value of Past Materials

First, let's look at what kinds of materials accumulate and how their retention rights differ. Different retention rights mean different limits on reuse.

Material Retention
Meeting notes (company material) Within company authority
Decision cards (company material) Within company authority
Quarterly retrospectives (personal + company) Personal copy allowed
Study notes (personal) Personal, permanent
Incident records (personal learning) Personal, permanent

Meeting notes and decision cards stay within company authority. Retrospectives can be kept as personal copies, and study notes and incident records are fully personal assets. Materials accumulated over many years are a major learning asset in their own right, but the boundary between company IP territory and personal territory must not blur. The clearer the boundary, the more comfortably you can reuse.


H.2 Separating Company IP from Personal Learning

The criterion for separation is "concrete or abstract." Concrete deliverables belong to the company; the thinking patterns that produced them belong to you. The key point is that both aspects come out of the same work.

Area Company IP Personal Learning
Decision content Company
Decision patterns (this kind of decision works in this kind of situation) Personal
Game data Company
Operating know-how (rulebook and tool operation) Personal
Code Company
Algorithms and structures Personal

"Which decision was made" is company IP, but the pattern — "in this kind of situation, this kind of decision tends to work" — is personal learning. The game data values themselves belong to the company, but the know-how gained from operating that data is yours. Keep the concrete materials at the company and take only the abstract patterns — that is the separation principle.


H.3 The Reuse Procedure

Turning the separation principle into actual work gives you the following five steps. Identify the materials, strip out the IP, extract the learning, generalize it, and apply it to the new project.

flowchart TD
    A["Identify past materials"] --> B["Separate company IP portions"]
    B --> C["Extract personal learning portions"]
    C --> D["Abstract and generalize"]
    D --> E["Apply to new project"]
    classDef pass fill:#dcfce7,stroke:#16a34a,color:#14532d;
    class E pass;

Always run this procedure only after checking company permissions and getting legal review. Even when the abstraction is thorough, if the starting point was company material, it is safer to have the procedural sign-off on record.


H.4 A Reuse Case — This Book

The nearest reuse case is this book itself. Much of the main text started from my past work and went through the procedure above to be generalized and anonymized.

Area Source Reuse
Layer-unified design (Part 6) My years of operation Personal learning → generalized
Meeting-notes system (Part 17) My Project A operation Company pattern → anonymized
Operating know-how (Part 24) Accumulated over years Personal learning → generalized
Appendix A inventory Company Project A Anonymized + partially reworked

The Layer design and the operating know-how generalize personal learning; the meeting-notes system and Appendix A anonymize company patterns. Every item passed company consent, and every piece of company IP was anonymized without exception. The book as a deliverable is itself a demonstration of the H.3 procedure.


H.5 The Five Pitfalls of Reuse

Done well, reuse is an asset; done badly, it is an incident. The five pitfalls below are spots people actually step on often, and each comes with a prescription.

H.5.1 Pitfall 1 — Skipping Company Approval

Using materials without the company's consent escalates into a dispute. The prescription is simple: get company consent before you use anything.

H.5.2 Pitfall 2 — Missed Anonymization

If a company name or a real name survives in even one place, it becomes an IP incident. The prescription is an automated grep check. Build a watchlist of company names, real names, and paths, and let the machine sweep them exhaustively.

H.5.3 Pitfall 3 — Applying the Past Unchanged

Old know-how used untouched will not fit the present. The prescription is to reconstruct it for the times: keep the principle, but update the tools and the context to today.

H.5.4 Pitfall 4 — Insufficient Abstraction

Moving only concrete cases makes them hard to apply in other environments. The prescription is to keep the abstract pattern and the concrete example side by side. The pattern carries the generality; the example carries the understanding.

H.5.5 Pitfall 5 — Skipping the Learning Itself

No matter how much material you have, if you never open it again, it might as well not exist. The prescription is a regular learning cycle. Like daily, weekly, and monthly retrospectives, build a cadence for meeting your materials again.


H.6 For the Reader — Reusing Your Own Materials

This principle is not mine alone. You can reuse the materials of your own career the same way. Below are recommended habits you can start today.

Recommendation Why
Retrospect on your own decisions every quarter Pattern discovery
Keep study notes separately Separation from company IP
State abstract patterns explicitly Future reuse becomes possible
Mentoring and external talks Pattern sharing
Books and blogs (after company consent) Learning lasts

Retrospect on your own decisions each quarter and patterns start to show; keep your study notes separate from company materials and you can pull them out later with peace of mind. Send those patterns out through mentoring, talks, and writing, and your learning lasts instead of being used once and lost. In the end, your own learning is your own asset.

Appendix I. Behavior Tree Editor Case Study (Advanced)

An advanced case study of the Behavior Tree editor covered in 7.2. The in-house development decision, implementation, and operating experience.


I.1 The Decision to Build In-House

Details behind the four decision rationales covered in 7.2.8.

Rationale Detail
Diff and git tracking are a must UE BTs are .uasset binaries, which makes change tracking hard. JSON allows text diffs
subtree references + impact tracking When running BTs for 100–300 NPCs, subtree-level impact analysis is decisive
Simulation verification BTs can be executed in isolation without a build
AI-assisted authoring LLMs generate and interpret JSON BTs naturally

I.2 Implementation Phases

[1. Design and Requirements Definition (1–2 weeks)]
   - Spell out the four requirements in writing
   - JSON schema design

[2. Runtime Implementation (3–4 weeks)]
   - JSON parser
   - BT execution engine
   - subtree reference resolution

[3. Editor Implementation (4–6 weeks)]
   - JSON editor (graphical)
   - subtree library UI
   - Impact analysis tool

[4. Simulator (2–3 weeks)]
   - Isolated BT execution
   - Statistics extraction

[5. AI Integration (2–3 weeks)]
   - LLM-assisted BT authoring
   - Context injection

[6. UE Integration (2–4 weeks)]
   - Conversion to and from UE BT
   - Build integration

About 4–6 months in total, with 1–2 developers.


I.3 Operational Design and Incident Preparedness

I.3.1 On the Operational Numbers

This editor is an in-house tool at the R&D stage; it never ran long enough or at a large enough scale to call the result "a year of measured operation." So, following this book's principle, I am not publishing made-up operational statistics here. The 100–300 NPC scale covered in I.1 is the design target that justified building in-house — not a measured result.

Of the limits the design nailed down, the one that actually went into code is the subtree reference depth limit of 5 (to prevent infinite recursion). Values like the number of BTs in operation or the number of simulation runs vary with project scale, so instead of writing made-up numbers, I recommend measuring them in your own environment.

I.3.2 Incidents We Prepared For and What We Learned

Incident Lesson
Infinite subtree references (recursion) Reference depth limit of 5
Simulation vs. live behavior divergence Monthly calibration of the simulation environment
Hallucinations in LLM-generated BTs Verification + stronger designer review
BT count explosion (unplanned) Quarterly cleanup cycle

I.4 Cost and ROI

The development and operating costs are estimates based on the schedule we set when building this tool, and the "effects" are not measurements but the directions the adoption was aiming for. Instead of made-up savings figures, I record only the direction.

Item Value Nature
Development cost Developers for 4–6 months Planned schedule (estimate)
Operating cost 1–2 developer-weeks per quarter (maintenance) Planned schedule (estimate)
Intended effect — operations staffing Compress the headcount dedicated to BTs when running enemy NPCs at scale Direction (unmeasured)
Intended effect — incidents Structurally reduce BT incidents such as subtree recursion and LLM hallucination Direction (unmeasured)

The payback period for the adoption cost depends on your project's NPC scale and labor costs, so I recommend measuring the items above in your own environment before judging. I make no claim like "it pays for itself within a year" — we do not have that number.


I.5 LLM Assistance Example

I.5.1 Writing a New Enemy NPC BT

Uses the prompt from 7.2.6. Below is an example that shows the output structure (a format example, not actual operational data).

{
  "bt_id": "bt_new_mage_v1",
  "category": "ranged_combatant",
  "tags": ["scholar_faction", "ranged", "magic"],
  "root": {
    "type": "selector",
    "children": [
      {
        "type": "sequence",
        "name": "low_hp_retreat",
        "children": [
          {"type": "condition", "fn": "hp_below", "param": 0.3},
          {"type": "subtree_ref", "id": "subtree_retreat_to_ally"}
        ]
      },
      {
        "type": "sequence",
        "name": "magic_attack",
        "children": [
          {"type": "condition", "fn": "enemy_in_range", "param": 15},
          {"type": "subtree_ref", "id": "subtree_magic_attack_pattern"}
        ]
      }
    ]
  }
}

After designer review: simulation → pass → applied to the build.


I.6 Notes for the Reader

Building your own Behavior Tree tooling tends to be justified at an operating scale of 100+ NPCs (a design judgment). Below that, UE's built-in BT is enough.

Alternative options: - BehaviorTree.CPP (open source, standard) - Behavior Designer (third-party commercial) - In-house development (maximum freedom, operational burden)

For selection criteria, see 7.2.8.

Appendix J. Abbreviations and Glossary

This appendix gathers in one place the abbreviations used in the main text and the terms specific to this book. The main text spells out each abbreviation once at its first occurrence, but if you are not reading in order, or forget one along the way, you can look it up here directly. Where one abbreviation carries different meanings depending on context, I listed both.

The glossary is grouped in this order: Team Size Tiers → Game Design Documents → Game Domain → Data and Operations → AI and Tools → UI and Accessibility Standards → Files and Formats. Think first about what kind of abbreviation you are looking for, and you can narrow down which group it belongs to. For example, DPS and TTK are in "Game Domain," KPI and DAU in "Data and Operations," and atom and JIT in "AI and Tools."

There are three notation rules. ① For general abbreviations, I listed the full name together with its meaning. ② For terms used only in this book, such as atom and Wrapper, the full-name column reads "(a term specific to this book)." ③ When one abbreviation has two meanings, such as PK (Player Kill in war contexts ↔ Primary Key in data contexts), the main text notes which sense applies at first occurrence, and this table lists both.

Team Size Tiers

This book does not pin team headcount to a specific number; it uses the following three tiers. The same technique calls for a different depth of adoption depending on team size.

Tier Headcount Description
Small up to \~10 From solo and hobbyist developers to single-digit teams. Adoption steps 1–2 are usually enough
Mid-size 10–50 The range my team — the source of this book's operational cases — belongs to. The scale where the cumulative effects of standardization and consistency automation become unmistakable
Large 100+ Multiple disciplines, multiple teams. The scale that justifies dedicated infrastructure and dedicated operations staff

Wherever the main text gives both the tier and the headcount range, as in "a mid-size (10–50 person) team," this table is the reference. Where the headcount itself carries the meaning — solo or one-person development — I write the exact number instead of the tier.

Game Design Documents

Abbreviation Full Name Meaning
GDD Game Design Document Game design document. The detailed spec that locks down systems, numbers, and behavior
CDD Concept Design Document Concept design document. The early-stage design doc (direction and concept) that precedes the GDD
TF TaskForce A dedicated team assembled temporarily for a short-term goal (e.g., a combat TF)
DD Design Director Design director. The lead role that oversees the game's overall design direction
RnD Research and Development Research and development. The phase or organization that explores prototypes and new techniques (e.g., procedural generation RnD)

Game Domain

Abbreviation Full Name Meaning
NPC Non-Player Character A character the player does not control
HUD Heads-Up Display Status information overlaid on the game screen (health, minimap, etc.)
DPS Damage Per Second Damage dealt per second
GCD Global Cooldown Global cooldown. A shared wait time that briefly locks all skills together after you use one
TTK Time To Kill The time it takes to kill a target
PK Player Kill (In war/PvP contexts) combat and kills between players
BT Behavior Tree Behavior tree. A structure that defines an NPC AI's behavior branches as a tree
FSM Finite State Machine Finite state machine. A model that defines behavior through states and transitions
PCG Procedural Content Generation Procedural content generation. Generating content automatically from rules and algorithms
VFX Visual Effects Visual effects
SFX Sound Effects Sound effects
VA Voice Actor Voice actor
RPG / MMORPG (Massively Multiplayer Online) Role-Playing Game Role-playing game / massively multiplayer online RPG
P2W / P2E Pay To Win / Play To Earn A structure where paying makes you stronger / a structure where playing earns you money
RMT Real Money Trading Trading in-game goods for real-world money

Data and Operations

Abbreviation Full Name Meaning
KPI Key Performance Indicator Key performance indicator
DAU Daily Active Users The number of daily active users
FK Foreign Key Foreign key. A column that points to another sheet's primary key
PK Primary Key (In data contexts) primary key. The column that uniquely identifies a row
ROI Return on Investment The return you get relative to what you invested
MECE Mutually Exclusive, Collectively Exhaustive Mutually exclusive, collectively exhaustive. A classification principle: no overlaps, no gaps
STT Speech-to-Text Converting speech to text
VBA Visual Basic for Applications The macro language built into Excel
SVN Subversion A file version control system
telemetry (measurement data) Play logs and metrics collected automatically from game builds and runs (inputs, combat, churn, etc.)

AI and Tools

Abbreviation Full Name Meaning
AI Artificial Intelligence Artificial intelligence
LLM Large Language Model Large language model (the foundation of ChatGPT, Claude, and the like)
JIT Just-In-Time Inserting something only at the moment it is needed (in this book, automatic injection of memory matched to the input)
MCP Model Context Protocol A standard for connecting AI tools to external services
API Application Programming Interface A convention for calls between programs
UE Unreal Engine Unreal Engine
atom (a term specific to this book) A decision/rule card codified as one decision = one file
Wrapper / Cascade / Junction (terms specific to this book) An entry point for frequently used tools / a tool that bundles multiple checks into a single run / a symbolic link that connects to the main copy
rg ripgrep A fast text search command (a CLI tool that replaces grep). Used for exhaustive searches across code and documents
ClickUp (a task/issue tracker) A cloud collaboration tool for managing tasks and schedules. JIRA, Redmine, and Linear are in the same category. Connected via MCP so the AI can query and update it

UI and Accessibility Standards

Abbreviation Full Name Meaning
UI / UX User Interface / User Experience User interface / user experience
WCAG Web Content Accessibility Guidelines Web accessibility guidelines (pass thresholds for contrast ratio, touch target size, and so on)
HIG (Apple) Human Interface Guidelines Apple's interface guidelines
SC Success Criterion The number of an individual WCAG pass criterion (e.g., SC 1.4.3)
pt / dp / px point / density-independent pixel / pixel Units of on-screen size

Files and Formats

Abbreviation Full Name Meaning
YAML YAML Ain't Markup Language A human-readable format for configuration and data
JSON JavaScript Object Notation A data interchange format
HTML / SVG HyperText Markup Language / Scalable Vector Graphics Web document format / vector graphics format
GLB GL Transmission Format (Binary) A binary file format for 3D models

For abbreviations like PK whose meaning splits by context, the main text notes which sense applies at first occurrence. If you lose track, come back to this table.

Appendix K. Porting to Other LLMs and Harnesses

Nearly every example and tool in this book assumes a single environment: Claude Code. So there is one objection that comes up almost without fail in sign-off meetings and external reviews: "Doesn't this tie us to one company's tool?" A head of design is reluctant to approve a decision that depends on a single vendor; a skeptic suspects that if the tool changes, this book's methods collapse wholesale; and those evaluating overseas publishing rights ask whether the book is still useful in a country where a different tool is the standard. The three phrasings differ, but the substance is the same: distrust of vendor lock-in — being locked into one tool.

The purpose of this appendix is to answer that distrust. The conclusion first: the skeleton of the work this book recommends is tool neutral. It is tied neither to a specific model name nor to a specific command-line tool. Claude Code was simply the vessel that implemented that skeleton most smoothly, and the same skeleton can be poured into other vessels. This appendix (1) shows in a table what the tool-independent skeleton is, (2) pairs each element of Claude Code with its counterpart in other environments, (3) sets a principle for checking what is current, on the premise that model generations keep changing, and (4) states honestly what you lose and what you keep when you move.


K.1 The Tool-Independent Skeleton

The way of working that runs through this entire book can be summarized as five pillars. None of the five is the feature name of a specific model or command-line tool; they are answers to the question of how people and AI, working together, can repeatedly produce results they can trust. That is why they remain when the tool changes.

Skeleton What it is Why it is tool neutral
Standard → template → verification gate Solidify agreed rules (the standard) into fill-in-the-blank forms (templates), and place an automatic checkpoint (the gate) that filters whether the result followed the rules Rules, forms, and checks are concepts any tool can express as text and scripts
atom = one decision, one file Write one decision in one small file, pull it out when needed, and when something changes, fix only that one cell Splitting decisions into small files requires nothing but a file system
JIT injection Pick out only the decisions the current conversation actually needs and feed them to the model just in time It is the principle of "inject only the necessary context"; only the injection method differs by tool
Retrospective loop Look back on the work daily, weekly, and monthly, and promote recurring patterns into rules for the next round of work The procedure of looking back and improving runs on habit and documents, not on a tool
Tool-borrowing boundary Borrow only the skeleton (algorithms, structure); leave the domain data behind (Appendix B) The judgment of what to take and what to leave is the same in any tool

The right-hand column of this table is the point. In all five pillars, not a single product name appears in the definition. What appears are universal concepts found in any work environment — rules, files, context, habits, boundaries. So the question "what do we do if we can no longer use Claude Code?" actually turns into a much easier question: "how do we implement these five concepts in another tool?" The answer is the next section.


K.2 Element Mapping Table (Claude Code → Other Environments)

Claude Code has concrete mechanisms that make implementing the skeleton above convenient: hooks (scripts that run automatically at specific points), MCP (a protocol that connects external tools and data to the model), settings files (permissions and environment configuration), slash commands (shortcuts that invoke a frequently used procedure in one line), and skills (reusable bundles of work). These names are specific to Claude Code, but their roles have counterparts in almost every other environment. The table below shows the pairs.

Claude Code ChatGPT (web/app) Cursor / Copilot Plain LLM API
hook (automatic execution at set points) manual pre/post-conversation procedures / custom GPT instructions pre/post editor tasks and pre-commit hooks pre- and post-call scripts wrapped around each call
MCP (external connection protocol) plugins / actions / code interpreter extensions / built-in tool calls function calling / hand-built API wrappers
settings file (permissions, environment) custom GPT settings screen / project settings .cursor and workspace settings files in-code config objects / .env and YAML config files
slash command (procedure shortcut) saved prompts / custom GPTs snippets / user-defined commands prompt template functions
skill (reusable work bundle) custom GPTs / prompt collections rule files + scripts modularized prompts and code functions
CLAUDE.md / memory custom instructions / memory feature project rule files (rules) system prompt + external memory store
atom file collection (tool-independent) Markdown files (tool-independent) Markdown in the repository (tool-independent) files and DB records

One thing becomes clear from the table. The further right you go — toward the plain LLM API — the more "things done for you automatically" turn into "things you have to build and wire in yourself." Automatic injection that took a single hook line in Claude Code becomes a pre-call script you write yourself against a plain API. The convenience of automation shrinks, but the skeleton itself carries over intact. Porting, in other words, is not "losing features" but "re-laying the conveniences with your own hands."

flowchart LR
    subgraph 도구중립["Tool-Neutral Skeleton (does not change)"]
        S[Standard→Template→Verification]
        A[atom: one decision per file]
        J[JIT injection]
        R[Retrospective loop]
    end
    subgraph 구현["Per-Environment Implementation (swappable)"]
        CC[Claude Code: hooks·MCP·settings]
        GPT[ChatGPT: plugins·custom GPTs]
        CUR[Cursor·Copilot: extensions·rule files]
        API[LLM API: function calling·pre-call scripts]
    end
    도구중립 --> CC
    도구중립 --> GPT
    도구중립 --> CUR
    도구중립 --> API
    classDef code fill:#dbeafe,stroke:#2563eb,color:#0b2545;
    classDef ai fill:#f3e8ff,stroke:#9333ea,color:#3b0764;
    classDef human fill:#fde68a,stroke:#b45309,color:#000;
    classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b;
    class J code;
    class CC,GPT,CUR,API ai;
    class R human;
    class A data;

This diagram is the one-page summary of the whole appendix. The upper box (the skeleton) keeps the same contents no matter which environment the arrows point to; only the lower box (the implementation) gets swapped to match the environment. When the words "vendor lock-in" come up at a sign-off meeting, open this one diagram and answer: what gets locked in is the lower box, not the upper one.


K.3 The Premise That Model Names Change

When discussing porting, the information that goes stale fastest is the model name. If I nailed the latest model name as of this writing into the text, that sentence would become wrong the moment the next generation shipped. So this book follows one principle from the start: never explain by leaning on a specific model's name or generation number; explain by leaning on the role the model plays (functions such as reasoning, summarization, or code generation).

What changes (do not nail down) What does not change (safe to lean on)
Model product names and generation numbers Role distinctions such as "a model that reasons well" or "a model that takes long context"
Specific figures for context limits The JIT principle: "there is a limit, so inject only the context you truly need"
Specific figures for price and speed The cost discipline: "run expensive work only on what has passed the gate"
How to switch a specific feature on and off "The role that feature plays," and the skeleton that can replace it

In practice, checking the latest models and features takes one line in any tool. In Claude Code, the /model command instantly shows the model in use and the available options, and tools like ChatGPT and Cursor show the same information in a settings screen or a model-picker dropdown. So if any sentence in this book seems to clash with a model name, the sentence is not wrong — the model is one generation on. As long as the role matches, the method applies as is. When a model name in the book looks unfamiliar, do not doubt the text; first run something like /model and check the current state of the tool in your hands.


K.4 What You Lose and What You Keep When Porting

Switching tools clearly loses you something. Hiding that fact would cost more trust than it saves, so I will state honestly what you lose first. But what you lose is almost entirely in the territory of "convenience," and what you keep is in the territory of "skeleton." That is, what you lose can be recovered by laying it down again, and what you keep was never tied to the tool in the first place.

Category Item Description
What you lose (convenience) The smoothness of automatic execution Automation that used to step in on its own, like hooks, must be rebuilt by hand as pre- and post-scripts
What you lose (convenience) One integrated screen Commands, tools, and files that gathered in a single flow may have to be spread across several tools
What you lose (convenience) Skills and commands that work immediately Slash commands and skills must be re-registered in the new tool's way
What you keep (skeleton) Standards, templates, verification gates Rules, forms, and checks are text and scripts, so they live anywhere as is
What you keep (skeleton) atoms, JIT, the retrospective loop They run on files and habits, so they survive a change of tools
What you keep (skeleton) The tool-borrowing boundary (Appendix B) The criteria for what to take and what to leave are independent of the environment

Compressed into one sentence, the table says this: what porting loses is the convenience of automation, which time can restore; what porting keeps is the skeleton of the work, which this book tried from the very beginning to keep outside any tool. So the most honest answer to "isn't this vendor lock-in?" is this: there is a part that gets tied down, but that part is a vessel you can swap out, and the real value — the contents — was never bound to any vessel to begin with. I hope this one appendix can stand in for that answer at the sign-off table.

Appendix L. Team Adoption TCO and Onboarding Worksheets

This appendix is a fill-in-the-blanks worksheet built to answer the question studio PDs (project directors) and CEOs ask: "When a one-person, six-month system scales to a mid-sized team, what do we use to estimate adoption effort, operating costs, accounts, and internal-network security — and how?" Where §19.3 in the main text (AI Adoption Strategy and Executive Buy-In) said "don't doctor the ROI," this appendix applies the same principle to the cost side of adoption. In other words, this appendix provides no numbers. Every field is blank; filling those fields is your team's measurement and estimation, and no one fills a field marked [pending finance confirmation] with an estimate before finance does.

Here is how to use this appendix. First, in L.1, get a picture of how TCO (total cost of ownership) breaks down into line items. Then print the five worksheets in L.2–L.6 — still blank — matched to your team-size row, and either measure the values yourself or hand them to finance or information security as one-line questions. Finally, run the self-checklist in L.7 to confirm no field is missing. The value of this appendix lies not in filled-in numbers but in turning the cost items that are easiest to forget into fields ahead of time.


L.1 TCO Is Not the License Fee

The trap PDs fall into most often is seeing adoption cost as nothing but "subscription × headcount." The real total cost of ownership is wider. It splits into one-time adoption effort (setup, standardization, onboarding) and recurring monthly operating costs (licenses, tokens, infrastructure, admin labor), with the less visible cost of security and account management layered on top.

flowchart TB
    TCO["Team adoption TCO"] --> A["One-time: adoption effort
(setup, standardization, onboarding)"] TCO --> B["Recurring: monthly operating cost
(licenses, tokens, infrastructure, admin)"] TCO --> C["Recurring: security and account management
(SSO, audits, key revocation)"] A --> A1["L.3 adoption effort worksheet"] B --> B1["L.5 monthly operating cost worksheet
[pending finance confirmation]"] C --> C1["L.4 internal network and security check"] A --> A2["L.6 onboarding time worksheet"] B --> A3["L.2 accounts and licenses worksheet"] classDef data fill:#e2e8f0,stroke:#64748b,color:#1e293b; class A1,A2,A3,B1,C1 data;

Of the three branches, the ones PDs most easily underestimate are the left (adoption effort) and the right (security and accounts). The license fee arrives written on a quote, but "the effort of organizing the standards and skills one person built by hand over six months into a form the team can share" and "the security review that decides how far external LLM calls are allowed on the internal network" appear on no quote — which is why they always overrun the schedule and the budget. The worksheets in this appendix exist to surface those invisible costs first, even if only as blank fields.

§19.3.6 in the main text said "costs get no absolute figures in this book — they are blanks to be filled by finance." This appendix lays out, item by item, where those blanks belong.


L.2 Accounts and Licenses Worksheet

Fill this table in first. Write down who uses which tool, and how those permissions get issued and revoked, alongside headcounts. Fill the headcount fields with your team's actual numbers; fill the price fields from quotes or public price lists. This book does not supply prices.

Item What to write Who fills it Your team's value
Seats per tool Number of accounts (seats) each tool needs Lead ______ seats
Permission tier distribution full / per-task cap / one-time contractor (Appendix C.1.2) Lead full __ / regular __ / contractor __
Per-seat price Monthly per-seat fee for each tool Finance/Purchasing ______ per seat per month
Shared key or not Team-shared API key vs. per-person keys InfoSec □ Shared □ Per-person
Issuance procedure Path and lead time for issuing accounts to new hires Lead ______
Revocation procedure Path for reclaiming keys/seats at departure or contract end InfoSec ______

There are two rules. First, do not issue standing seats to contractors and short-term staff — open and revoke access per task (Appendix C.1.2). Second, if the revocation procedure field is blank, do not start issuing. The most common incident is a departed employee's account going unrevoked, leaking cost and key exposure at the same time — so design revocation before issuance.


L.3 Adoption Effort Worksheet by Team Size

This table estimates, by team size, the one-time effort that grows when "one person, six months" scales to a team. Fill the effort fields in person-days (the amount of work one person does in one day), measured or estimated by your own team. This book provides no person-day figures — they vary too widely with the team's skill level and how organized its existing standards already are.

Adoption effort item 1–3 people 4–10 people 11–30 people 31–50 people Measured/estimated by
Environment install and setup (tools, hooks, permissions) ___ person-days ___ person-days ___ person-days ___ person-days Lead/Infra
Turning one person's assets into team-shared form (organizing skills, standards, atoms) ___ person-days ___ person-days ___ person-days ___ person-days Lead
Establishing team standards (naming, frontmatter, rulebook — Appendix D) ___ person-days ___ person-days ___ person-days ___ person-days Lead
Building verification gates (lint and rulebook automation) ___ person-days ___ person-days ___ person-days ___ person-days QA/Lead
Producing onboarding materials (tied to L.6) ___ person-days ___ person-days ___ person-days ___ person-days Lead
Total (one-time adoption effort) ___ person-days ___ person-days ___ person-days ___ person-days

The field most easily skipped in this table is the second row. The assets one person piled up over six months — in their head and in personal folders — take separate, dedicated effort for someone to pull out, organize, and turn into documents before a team can share them. Budget this effort at "0" and the adoption schedule slips, without exception. The table's shape also shows in advance that as headcount grows, standards-setting and verification-gate effort climbs more steeply than install effort — more people means more standards that have to be agreed on.

Follow the staged adoption in §19.3.1 (conservative → progressive) and you don't have to spend this effort in a single quarter; you can spread it out, starting from the Stage 1 (context injection) pilot. Don't try to get the table's total approved in one shot — peel off just the Stage 1 effort and get that approved first. That is the realistic move.


L.4 Internal Network and Security Check Worksheet

This is the area PDs and CEOs fear most directly. It checks, item by item, what leaves for external LLMs and how far external calls are allowed from the internal network. This table is a checklist that sorts pass from hold (tied to Appendix C.6, Security); if even one item is undecided, adoption in that scope goes on hold.

Check item Pass bar Owner Status
Scope of data sent to external LLMs Sensitive data via placeholders / self-hosting (C.6) InfoSec □ Pass □ Hold
Payment and personal data transmission Prohibition codified, no exceptions InfoSec □ Pass □ Hold
Internal-network outbound call policy Allowed domains, proxy, log retention period defined Infra □ Pass □ Hold
Whether self-hosting is needed Decide whether core IP runs on a self-hosted model CEO/InfoSec □ Decided □ Undecided
Key exposure incident response Immediate rotation + usage history review path (C.7) InfoSec □ Pass □ Hold
Audit logs Who called what, and when — recorded and retained Infra □ Pass □ Hold
Company IP leak screening Advance screening such as a grep watchlist (Appendix B.6) Lead □ Pass □ Hold
Contractor access isolation Contractor accounts blocked from core assets, isolated per task InfoSec □ Pass □ Hold

The field where cost diverges most in this table is the fourth row (whether self-hosting is needed). Decide that core IP can never be sent to an external LLM, and the entire cost of self-hosted infrastructure lands on the operating costs in L.5. That is why this decision belongs not to the lead but to the CEO and information security together, and until it is made, the infrastructure field in L.5 cannot be finalized. The two worksheets connect through this one field.


L.5 Monthly Operating Cost Worksheet [Pending Finance Confirmation]

This table breaks the recurring monthly costs into line items. Every amount field in this table is blank, and no one fills a field marked [pending finance confirmation] with an estimate before finance does. Token prices, subscriptions, and infrastructure fees shift every month with models, call volume, and contracts, so this book writes no absolute figures.

Operating cost item How it's calculated Who fills it Monthly amount
Licenses and subscriptions Seat count × per-seat price (L.2) Finance [pending finance confirmation]
LLM token costs Call volume × token price, sum of per-tool caps Finance [pending finance confirmation]
Infrastructure (if self-hosting) Servers, GPUs, storage per the L.4 decision Finance/Infra [pending finance confirmation]
Backup and sync Repository and backup storage (Appendix C.5) Finance [pending finance confirmation]
Operations admin labor Hours spent managing tools, keys, and logs, converted to cost Lead/Finance [pending finance confirmation]
Monthly total Sum of the items above Finance [pending finance confirmation]

This table has exactly one rule: a blank stays blank. Recall the failure in §19.3.2, where the AI fabricated a plausible $4,500 into the operating-cost field — human or AI, the moment someone fills this field with an estimate, the report collapses at the first question. The real device that controls cost is not an amount but a structure where each tool carries a monthly cap and overruns get reported automatically (§19.3.6). What you show executives at approval time is not filled-in amounts but that structure — "caps are in place and overruns get reported" — plus the list of blanks finance will fill.

The fifth row (operations admin labor) is the one most often dropped. A tool is not done once installed; it eats someone's hours every month — revoking keys, reading logs, adjusting caps. Leave this field at 0 and that work hides as the lead's invisible overtime.


L.6 Onboarding Time Worksheet

This table estimates, stage by stage, how long it takes one new member to start pulling their weight on top of the system. The most accurate way to fill the time fields is to actually run one onboarding on your team and measure it (the same method as the baseline measurement recipe in §19.3.7). Until measured, leave them blank.

Onboarding stage What happens Measured time Notes
Environment setup Through tool, hook, and account setup ___ hours Tied to the L.2 issuance procedure
Learning the standards Naming, frontmatter, rulebook (Appendix D) ___ hours Shorter if materials exist
First task (conservative) First deliverable via context injection, passing review ___ hours Stage 1 of §19.3.1
Adapting to verification gates Working within the lint and rulebook gates ___ hours
Reaching independent work Can work and make accept/reject calls unsupervised ___ days The bar for onboarding completion

Fill in this table and it becomes clear why the "producing onboarding materials" field in the adoption effort worksheet (L.3) matters. The better organized the onboarding materials, the shorter the second and third rows — and the more new members you add, the more that saving compounds. Producing onboarding materials is a one-time effort, but the payback repeats once per member. What §19.3.3 in the main text meant by "221 JIT auto-injections — new members work on top of the same rules" shows up in this table as a shorter third row.

The last row (reaching independent work) is the real completion bar for onboarding. Mistake "environment setup finished" for "onboarding finished," and the supervision cost keeps piling onto the lead. The bar has to be "makes accept/reject calls without supervision."


L.7 Pre-Adoption Self-Checklist

Finally, these are the items to pass on your own before taking the worksheets to executives. In the same spirit as Appendix B.6 (the pre-adoption check), if even one item is blank, postpone the approval request and fill that field first.

Check item Pass bar
Is the account revocation procedure defined? The revocation procedure field in L.2 is not blank
Have all security checks passed or been decided? Zero □ Hold / □ Undecided entries in L.4
Have the operating-cost blanks gone to finance? The L.5 [pending finance confirmation] items sent out as questions
Is the adoption effort split into stages? Approval starts from the Stage 1 effort, not the L.3 total
Is the onboarding completion bar "independent work"? Completion judged by the last row of L.6
Are no estimates written as assertions? Every estimated field labeled with "estimate + sample size"

Please don't read this table as five boxes to pass — read it as six locks. Scaling a one-person system to a team is absolutely possible, but the cost of that scaling is not the license fee; the full picture appears only once you have honestly filled in these six blanks. And don't ask the AI to fill in any of them — as in §19.3.2, AI fills blanks with plausible numbers. The AI's seat ends at taking the values you measured and shaping them into the sentences on the approval slide.

Appendix M. Dimension Vectors and Embeddings — An Intuition for Game Designers

In five places in this book — §8.2.7 (economy), §5.4 (voice), §6.3 (personas), §7.3 (patterns), and §13.3 (data) — expressions like "compress into a dimension vector," "embedding," and "close in vector space" appear as signposts. Even without a machine learning background, the concept is fully graspable with a game designer's instincts alone. Lay the intuition down once, and all five of those places read as one and the same picture.

But let me nail one thing down first. The intuition behind the concept is easy; the conditions for applying it are heavy. This idea is not an entry point — it is the application at the farthest end, one that can stand only on the foundation this book has been stacking up all along: the verification gates of conservative application, data and telemetry infrastructure, per-domain checkers like voice_lint and consistency checks, and Layer unification. Draw the coordinates first without that foundation, and — as M.4 shows — the map becomes an illusion that has neatly compressed even the error by which it diverges from the game. Here is the real reason all five places say "still premature" — not because the idea is hard, but because the supporting foundation comes first. This appendix is the map a team that already has that foundation opens to gauge its next step, not a primer for taking the first one.

M.1 One Line — Features Become Coordinates, Similar Things Become Nearby Points

Turning any object's features into a list of numbers and placing it as a single point on a "map" — that is an embedding. That list of numbers is the dimension vector. There is only one promise — you build the map so that the more similar two objects are, the closer their points end up.

For example, place an NPC at the coordinates of features like (formality of speech, emotional expressiveness, vocabulary difficulty, ...), and NPCs with similar voices gather close together on the map. For cooking recipes, use the ingredient composition as the coordinates and similar dishes gather close together — this is exactly what Epicure, cited as a lead in §8.2.7, did.

Features as coordinates — 3D is a metaphor; reality is hundreds of dimensions Feature 1 Feature 2 Feature 3 … (hundreds) Back cluster (fainter when farther) Empty region no such combination yet Front cluster (darker when closer) Join two points: what lies between = interpolation (a middle variation) Nearby points = similar · empty region = no such combination yet · between two points = interpolation

M.2 Once the Map Is Drawn, Three Things Become Visible as Distance and Position

  1. Nearby points = similar things. When the points clump into a single spot, that is the signal that "diversity has died" (you read §5.4's voice convergence and §6.3's persona staleness not as an impression but as point density).
  2. An empty region = a combination that doesn't exist yet. A spot with no points at all is a design blank that says "this design doesn't exist yet" (§7.3's empty pattern regions, §13.3's emergence of new topics).
  3. Between two points = a middle variation. Connect two points and take what lies between them, and you get a "middle" — this is interpolation. Join "rice" and "curry leaf" and you get the flavor in between; join two NPCs and you get the persona in between.

M.3 Three Terms That Come Up Often

Similar-looking but opposite — do not confuse this with AHP. AHP (Analytic Hierarchy Process, Saaty), used for multi-criteria decision making, can look like a cousin because it too turns qualitative judgments into vectors — it derives priority weights (the principal eigenvector) from pairwise comparisons that weigh criteria two at a time. But the direction is reversed. AHP is top-down decision making in which a person predefines the criteria and the hierarchy and assigns weights within them; the embedding here is bottom-up discovery in which clusters surface from the data on their own, with no human definitions. The limitation §13.3 tries to break through — "segments a person defined in advance" — is exactly where AHP starts. The two do not sit in the same spot; they sit on opposite sides.

M.4 Why This Book Writes "Still Premature" — The Price of Compression

The map is powerful, but it is not free.

That is why this book keeps dimension vectors as a signpost, not a prescription — territory for a team with verification (telemetry and simulation) laid down solid to look into a few years from now. The concrete per-domain leads are scattered across §8.2.7 (economy), §5.4 (voice), §6.3 (personas), §7.3 (patterns), and §13.3 (data), and you can read them all on this appendix's single map.

Appendix N. A 15-Week Course Schedule and Difficulty Guide

This appendix is for anyone who wants to build a semester-long course on top of this book — instructors at universities, colleges, and academies, in-house training leads, study-group organizers. Splitting a single volume of nearly 1,000 pages across a semester is more daunting than it looks. Which parts go into which weeks, how the Try It Yourself sections in the body become assignments, what criteria to grade submissions by — when those three questions stall, even a good book is hard to adopt as a textbook. This appendix turns those three things into tools you can copy as is.

Here is how to use it. First, read the 15-week schedule in N.1 against your own academic calendar (16-week and intensive-term variants are kept separately in N.2), then gauge your students' level with the difficulty badges and prerequisite table in N.3. After that, copy the grading rubric in N.4 and swap in the items for your own assignment. Every table is built so you can print it as is and paste it into your syllabus.

One thing worth saying up front. Every chapter in this book ends with a Try It Yourself. The body's goal was never a chapter you read and close — it was a chapter that gets your hands moving today. In a course, that very Try It Yourself becomes the primary raw material for assignments. That is why the schedule in this appendix also notes how each Try It Yourself in the body is carried over into a weekly assignment.


N.1 The Standard 15-Week Schedule

This is the standard schedule, built for the most common 15-week semester (one 3-hour session per week). It does not cover all 24 parts of this book in one semester — cram them in and nothing stays in anyone's hands. Instead, the structure I chose is this: lay the foundation (Parts 1 and 2) solidly, go deep on five or six representative domains, and close with only the essentials of process and operations. Parts left uncovered are marked as Further Reading, so students with the interest can open them on their own.

Every learning objective is written as a verb of what students will be able to do — not "knows" but "builds, verifies, chooses." This whole book repeats a single sentence — the AI proposes candidates and a human filters them — so the verbs in the objectives follow that division of labor.

Week Parts and Chapters Learning Objectives (What Students Can Do Afterward) Try It Yourself Converted into the Assignment
1 1.0 Before You Start + Part 1 (Introduction) Explain terminals, accounts, and pricing, install the AI tool on their own PC, and open a first session The 1.0 setup Try It Yourself — submit an installation screenshot plus the first prompt and output
2 Part 2 (Information Architecture) Turn documents into data with YAML frontmatter and design folder and naming conventions The 2.1 frontmatter Try It Yourself — add frontmatter to three of their own documents
3 Part 3 (Systems Design) Define a data sheet's $schema first, following the schema-first principle The 3.2 schema Try It Yourself — write the spec for one mini sheet
4 Part 10 (QA and Integrity) Follow along and build a tool that checks FK integrity across 30 sheets in code The 10.1 integrity check Try It Yourself — the core assignment, graded with the N.4 rubric
5 Part 4 (Combat) + Part 8 (Balance) Decompose combat numbers into Layers and keep a deterministic balance formula as a rulebook The 8.1 balance formula Try It Yourself — one damage formula plus a simulation
6 Part 5 (Narrative) Build an NPC dialogue voice_profile and catch tone drift with voice_lint The 5.2 voice_profile Try It Yourself — a voice profile for one character
7 Part 6 (Content) + Part 7 (Level) Distinguish the two axes of procedural generation (rules vs. AI) and mass-produce and review content candidates The 6.2 generator Try It Yourself — generate 10 content candidates plus a review log
8 Midterm Check-in and Presentations Integrate the Weeks 1–7 assignments and demo them as their own mini project Midterm presentation (an integrated demo of the Weeks 3–6 deliverables)
9 Part 9 (UX/UI) + Part 14 (Mobile) Run a HUD through lint to catch gaze drift and contrast shortfalls, and compress a PC HUD for mobile The 9.1 HUD lint Try It Yourself — a lint report for one screen
10 Part 16 (Communicator) + Part 17 (Meeting Notes) Canonize only the decisions from an isolated workspace, and structure meeting notes The 17.x meeting notes Try It Yourself — structure one real meeting recording
11 Part 18 (Decision-Making) + Part 19 (Team Lead) Leave decisions as traceable cards and turn the vision into a scorecard for decisions The 18.1 decision tracking Try It Yourself — write three decision cards
12 Part 20 (Collaboration Memory) + Part 21 (Self-Improvement) Run collaboration context as memory and turn retrospectives into a self-improving loop The Part 21 retrospective Try It Yourself — one weekly retrospective plus one extracted rule
13 Part 22 (Governance) Inspect the boundaries of prompts, hallucination, cost, legal, and ethics, and set the rules The 22.1 prompt Try It Yourself — one work order plus a hallucination-check procedure
14 Part 23 (Personal Development) + Part 24 (Operations Deep Dive) Port the tools to a solo scale-down, and verify integrity, links, and staleness in code The 24.1 verification Try It Yourself — one verification script for their own project
15 Final Project Presentations and Evaluation Design, demo, and verify one workflow of their own that spans the whole semester Final presentation (evaluated with the extended N.4 rubric)

Further Reading (not covered in the course; self-study recommended): Part 11 (Characters, Pets, and Mounts), Part 12 (Art Direction), Part 13 (Data and KPIs), Part 15 (Live Ops). These four parts are strongly domain-specific, so I left them for students to open according to their own field. Appendix F (the case index) works as a guide: students can start from the cases closest to their own environment and trace backward from there.

The flow of the schedule at a glance: foundations → domain deep dives → midterm integration → process and operations → final integration — a structure with two peaks (midterm and final).

flowchart LR
    subgraph A["Foundations (Weeks 1–3)"]
        W1["Week 1
Setup · Intro"] --> W2["Week 2
Information Architecture"] --> W3["Week 3
Systems · Schema"] end subgraph B["Domain Deep Dives (Weeks 4–7)"] W4["Week 4
QA · Integrity"] --> W5["Week 5
Combat · Balance"] --> W6["Week 6
Narrative"] --> W7["Week 7
Content · Level"] end M1{{"Week 8
Midterm Presentations"}} subgraph C["Process · Operations (Weeks 9–14)"] W9["Week 9
UX · Mobile"] --> W10["Week 10
Communication"] --> W11["Week 11
Decisions · Leads"] --> W12["Week 12
Collaboration · Retrospectives"] --> W13["Week 13
Governance"] --> W14["Week 14
Personal Dev · Verification"] end M2{{"Week 15
Final Presentations"}} A --> B --> M1 --> C --> M2 classDef human fill:#fde68a,stroke:#b45309,color:#000; class M1,M2 human;

N.2 Semester-Length Variants (16 Weeks / 8-Week Intensive Term)

Semester lengths differ from school to school. Beyond the standard 15 weeks, here are adjustments for the two variants you will meet most often. I recommend keeping the core assignment (the Week 4 integrity check) and the two presentation peaks in any variant — they are where this book's honesty principle ("show the structure, not the effect") comes through most clearly.

Semester Format How to Adjust
16-week The standard 15 weeks, plus a make-up and reassessment week in Week 16. A chance to resubmit the final project, or a special session on one of the four Further Reading parts chosen by student vote
8-week intensive term (twice weekly or condensed) Week 1 (setup and intro) → Week 2 (information and schema) → Week 3 (integrity, the core assignment) → Week 4 (combat, balance, and narrative bundled) → Week 5 midterm presentations → Week 6 (meetings, decisions, and collaboration) → Week 7 (governance and verification) → Week 8 final presentations. Cut the domains down to three representatives, and absorb the Try It Yourself sections into in-class exercises
Flipped classroom Move the reading to pre-class assignments and give class time entirely to Try It Yourself exercises and rubric-based peer review. The code in this book runs as is, with no external dependencies, which suits exercise-centered teaching

N.3 Chapter Difficulty Badges and Prerequisites

Even within the same book, chapters demand different background knowledge. Some chapters can be followed by a first-year student who has never opened a terminal; others need database key concepts or basic statistics to digest fully. I organized them into three badge levels, for pacing the course to your students' level or pointing them to prerequisite courses.

The badges mean the following.

Badge Level Meaning
🟢 Intro Intro Non-majors and first-years can follow. Copying and running the code is enough
🟡 Practitioner Practitioner Students should be able to read the code and adapt it to their own data. Familiarity with working design practice recommended
🔴 Advanced Advanced Designing and extending algorithms and structures. Hard to digest without the prerequisites

Badges and prerequisites for each week's core parts are below. "Prerequisites" are background that helps students follow the week comfortably — lacking them does not bar anyone from taking the course.

Week Core Parts Badge Prerequisites
1 1.0 and the Part 1 introduction 🟢 Intro None (assumes first contact with a terminal)
2 Part 2 Information Architecture 🟢 Intro Using a text editor
3 Part 3 Systems and Schema 🟡 Practitioner Table/spreadsheet basics, the concept of data types
4 Part 10 Integrity Verification 🔴 Advanced Python basics (functions, loops), the concept of relational keys (FK)
5 Parts 4 and 8 Combat and Balance 🟡 Practitioner Arithmetic formulas, spreadsheet calculation (Excel functions)
6 Part 5 Narrative 🟢 Intro A feel for character and scenario writing
7 Parts 6 and 7 Content and Level 🟡 Practitioner Procedural generation concepts (recommended), a feel for coordinates and grids
9 Parts 9 and 14 UX and Mobile 🟡 Practitioner Screen layout and resolution concepts
10 Parts 16 and 17 Communication 🟢 Intro None (collaboration experience helps)
11 Parts 18 and 19 Decisions and Leads 🟡 Practitioner Teamwork and project management experience (recommended)
12 Parts 20 and 21 Collaboration and Retrospectives 🟡 Practitioner Completion of Week 2 (information architecture)
13 Part 22 Governance 🟡 Practitioner Basic statistics (means and distributions, for the hallucination-detection context), copyright basics
14 Parts 23 and 24 Personal and Operations 🔴 Advanced Python basics, git basics, completion of Week 4 (integrity)

Prerequisite one-liner (for the syllabus): "An introductory Python course or equivalent programming basics is recommended but not required. The advanced chapters in Weeks 4 and 14 assume Python functions and loops; students without that background can fully keep up via the Weeks 1–3 intro track, with assignments run on a separate track."

Tips for different class compositions:


N.4 Sample Grading Rubric — The Integrity Checker Try It Yourself (Week 4 Core Assignment)

Without a rubric, Try It Yourself submissions tend to get graded on a binary: it ran or it did not. That erases from the assessment the thing this book values most — the process of reviewing and rejecting AI output. So here is a rubric that grades the process as well as the result, using the Week 4 core assignment (the 10.1 integrity check atom Try It Yourself) as the example. For other weeks' assignments, swap the item names and use it as is.

Assignment definition: Working with an AI, build a tool that checks foreign key (FK) integrity across several data sheets the student made (or was given), and demonstrate that the tool catches errors planted on purpose. Submissions: ① the tool's code, ② the check run's results (a pass/fail report), ③ the full text of the prompts given to the AI, with a record of which outputs were rejected or revised.

The rubric has four items at 25 points each (100 points total). The key point: separate from "the tool runs" (item 2), how the student handled the AI (items 3 and 4) carries half the weight.

# Criterion Points Poor (0–12) Fair (13–19) Excellent (20–25)
1 Integrity rule definition — is it clear which FK relations are checked and why 25 Target relations unclear or arbitrary Major FK relations identified, but the rationale is thin Relations between sheets defined with diagrams and rationale, and check priorities explained
2 Tool behavior and error detection — does it actually catch the planted errors 25 Does not run, or misses obvious errors Catches most errors, with some misses or false positives Catches every planted error, no false positives, and prints a report a human can read
3 Transparency of the AI process — are the full prompts and outputs recorded reproducibly 25 No prompt/output record, or results only Prompts present, but the rejection/revision process is missing Full prompts, raw outputs, and the reject-and-redirect process kept in chronological order
4 Review and rejection judgment — what in the AI output was rejected or revised, and why 25 Output accepted as is (no trace of review) Some revisions, but the rationale is weak Errors, hallucinations, and overengineering identified and rejected, with the rationale explained in the student's own words

Grading note: Items 3 and 4 (50 points together) are the spine of this rubric. Even if the tool runs perfectly (full marks on item 2), a student who accepted the AI output uncritically (poor on item 4) has missed this assignment's learning objective — the human keeps the reviewer's seat. Conversely, a student whose tool is somewhat incomplete but whose reject-and-redirect process is solid can score high. The book's principle — evaluate the structure (how it was handled), not the effect (the fact that it ran) — applies to grading just the same.

For the final project, I recommend an extended rubric: the four items above plus ⑤ workflow generalization (explaining the port to the student's own field) — five items at 20 points each. Can the student carry the tools the semester covered into their own project — that is the question this book asks at the end, and it is enough for the course's final assessment to ask the same one.


N.5 The Course on One Page

Finally, this appendix reduced to a single page.

This schedule is a starting point, not the answer. Move the weeks and change the assignments to fit your students' level and your academic calendar. Feeding this entire book to an AI tool and asking, "rebuild this schedule for my 16-week course and my students' level" — fittingly, as the fastest way to use this book — is an open path too.


Colophon

AI Workflow for Game Designers

No Fabricated Numbers — A Six-Month Field Manual for Claude Code, Prompts, Validation, and Production Memory

This is the English edition of the Korean original, translated from Korean with the book's own AI workflow and reviewed by the author. This web edition carries no separate ISBN.

Original title 게임 기획 실무에서 바로 쓰는 AI·클로드 코드 활용법
Author Minsoo Lee (이민수)
Original publisher BOOKK Co., Ltd. (Korean print edition)
Original published June 11, 2026
Original ISBN 979-11-12-21479-9 (Korean print edition)
Korean source https://github.com/eremes81/game-design-ai-practice

ⓒ Minsoo Lee 2026

This book is released under CC BY-NC-SA 4.0. Noncommercial sharing and translation are allowed with credit to the original author (Minsoo Lee · 이민수) and the source; commercial use requires the author's separate permission.