One discipline · three doors contextkeeping.com machinereadyknowledge.com answerecon.com

The vocabulary layer — an essay on terminology governance

The language of knowledge work is failing its own test.

I said "tribal knowledge" in meetings for years. It used to be the knowledge industry's own term for the thing we exist to fix — the answers that live in people's heads and walk out the door at 5pm. I stopped using it, and the reason I stopped is the subject of this essay, because it turned out to be a better reason than the one I expected.

The expected reason is the inclusive language one, and it's real: the term borrows Indigenous identity as a metaphor for "unmanaged," and the people it borrows from have asked us to knock it off. The knowledge-management discipline has, mostly, complied; you'll find institutional knowledge in the recent literature where the older decks said tribal. But if that were the whole argument, this would be an etiquette essay, and etiquette arguments have been losing recently. While every organization I've ever worked in has adopted inclusive language initiatives and driven them to success, there's still so much more work to do, and inclusivity needs a more tactile and quantifiable friend.

Here's the argument that wins: "tribal knowledge" is a bad retrieval key. It's imprecise — it names two different things (knowledge nobody wrote down, and skill that resists writing down; institutional and tacit knowledge respectively, which is why the replacement is two terms, not one). It's idiomatic, so it means nothing reliable to a non-native English speaker. And it's exactly the kind of phrase that a machine, assembling an answer from your corpus, matches weakly against everything... and strongly against nothing.

That's the essay in one specific term. Now the general case.

Terms are retrieval keys

Run this test on your own knowledge base: pick a concept your organization touches every day, and count its names. In one audit I use as a working example, the company's own product appeared as "the endpoint agent," "the client," "the sensor," and "the collector" — with "the client" ambiguous against client, the customer in thirty-one places. Nobody decided this. Terms drift in from engineering slang, sales habits, and acquisition vocabulary, and nothing checks them on the way into the corpus.

For a human searching, that concept is now four searches, three of which return partial results. For a retrieval system, it's four weak clusters instead of one strong one. Every downstream AI ambition — the chatbot, the agent-assist panel, the customer's own AI querying your docs — inherits the split. You will spend real money on retrieval engineering to compensate for a problem that is, at root, editorial: the concept never got one name.

Librarians solved this a century ago and called it controlled vocabulary — the foundation for taxonomies and ontologies. It is the least glamorous idea in information science and, right now, one of the most valuable: one concept, one canonical term; synonyms map to it, never compete with it. Machine-ready knowledge already demands self-contained sections, explicit applicability, stable identifiers, and freshness signals. Vocabulary is the property under all four — a stable identifier for ideas. An article can pass every structural test and still be unfindable because it calls the thing by the wrong name.

Idioms fail twice

The second family of failures isn't drift; it's color. Open the kimono. Drink the Kool-Aid. Low man on the totem pole. Mexican standoff. Each of these is doing two jobs badly. A non-native reader — which is to say, a large segment of every global workforce and customer base — gets nothing from them. And a retrieval system gets less than nothing: an idiom doesn't match the plain-language query a person actually types when they have the problem the idiom is decorating.

Some of these phrases carry a third cost, and I'm not going to pretend otherwise: an ethnicity attached to dysfunction, a sacred practice borrowed as office furniture, a slur doing duty as a verb. You don't need two arguments — every one of these terms fails on precision alone, and the replacement is clearer in every case I've ever tested. But I've watched organizations tie themselves in knots trying to make this case without mentioning the freight, and the evasion reads as exactly what it is. Say both things plainly: the term is imprecise at best, and racist at worst. Then, move on to the fix. The audiences that matter can tell the difference between honesty and theater.

The paradox that raises the stakes

There's a compounding effect worth naming. As AI answers get more fluent, each remaining defect costs more — a polished answer built on an ambiguous term doesn't look wrong, it looks authoritative. The old, messy knowledge base advertised its own unreliability. The new one launders vocabulary drift into confident prose with your logo on it. The better your delivery gets, the more your inputs' precision matters. Language governance is input hygiene for machine answers.

What governance actually looks like

Here is what I mean by governing vocabulary, and — just as important — what I don't.

An inventory, in tiers. The industry has already done the consensus work on a surprising amount of this: the Inclusive Naming Initiative (that's Cisco, IBM, Red Hat, the Linux Foundation), the ACM, Microsoft's and Google's style guides, the Linux kernel, 3GPP. Tier 1 is the list with that consensus behind it — master/slave, whitelist/blacklist, grandfathered, sanity check, man-hours, and yes, tribal knowledge. Tier 2 is judgment calls — war room, postmortem, kill switch — where the precision case is real and the decision is legitimately yours.

A standard, ratified policy. Two pages: canonical terms, deprecated terms, a precision rule (a replacement must be more precise than what it replaces, or keep looking — this is never about euphemism), and named owners. The working example I keep returning to made three Tier-2 decisions in one afternoon: adopted incident room (the precision case won — the room handles incidents, name it so), kept postmortem (industry-standard among its actual participants), and split kill switch by audience (precise internally, "emergency stop" in customer copy). One adopted, one kept, one split. That's what governance looks like: decisions, made once, logged, defensible. Not vibes, and not a purge.

Enforcement where review already happens. Three line-items in the existing content gate, and a grep list generated from the standard's deprecated-terms table — a linter for language, run in CI at a tooling cost of zero. No language police, no new meetings. Speech gets coached, artifacts get fixed at the gate, and the standard's owner handles the contested calls so nobody else has to relitigate them.

A thirty-day proof. Twenty client-facing assets, two independent reviewers, thirty days. Three numbers come out: findings per asset, the count of charged terms on public surfaces (each one a screenshot away from being someone else's post about you), and — the one I care most about — the consistency rate. That is, of concepts appearing in three or more assets, how many carry one canonical name. That last number is your retrieval health, expressed as an editorial statistic. Re-audit in two quarters; the delta is the program.

The discipline, applied to its own words

I find it fitting that the knowledge discipline's most famous piece of vocabulary drift is its own term for undocumented knowledge. We of all people should have named that concept precisely — naming concepts precisely is the job. Knowledge Management, now Context Management, and in turn Contextkeeping treats language the way it treats everything else in the corpus: as an asset with an owner, governed at a gate, measured by whether the next reader — human or machine — gets the answer.

Your vocabulary is either an asset or a liability. The only way it becomes an asset is if somebody owns it.

Start with the audit. It's twenty assets and a month, and the findings report will make the rest of the argument for you — in your own corpus's words.

The twelve swaps and the four-move program fit on one page. Free, no email required. The full machinery — the 40-entry inventory, the standard template, the gate line-items, the audit blueprint, and the champion deck — is the Terminology Governance Pack; the audit, run for you with findings in 30 days, is an AnswerEcon engagement.

KCS® is a service mark of the Consortium for Service Innovation. This essay reflects an independent, KCS-aligned methodology and is not affiliated with or endorsed by the Consortium. Industry consensus sources cited: the Inclusive Naming Initiative, the ACM's "Words Matter," the Microsoft and Google style guides, the Linux kernel coding-style document, and 3GPP's specification-language initiative.