Payments: Sep 14 Invite coupons available · Workspace open
KBMILL Notes Enter the plant

Teaching

Notes. Not a blog.

Plain teaching for people who tried “just ask the PDF” and got confident wrong answers — and for anyone who needs a clear next step.

What we do: KBMill is a plant behind a public hopper — not a PDF parser with philosophy. It turns messy documents into a package you keep and point your own AI at — not a chatbot. Weak or unreadable material is set aside and listed, not hidden. You pay only if we deliver. What actually runs: behind the hopper.

Look, don’t trust us. Open sample packages on the public shelf, re-run the retrieval checks, and read the package contract. Machine brief: /llms · glossary: glossary. Door facts (limits, pay, purge): FAQ.

Questions we answer

These are questions we believe matter for the problems we solve. Each links to a Note (or FAQ) with our read and what we built. Structured as an FAQ index for humans, search engines, and models.

  1. What actually runs behind the hopper? Infrastructure-class plant under the hopper — not a parser with philosophy. Scales by mirroring plants; minimal pain next to your stack (keep the ZIP, keep your RAG, pay on success). Worth diligence if document prep is on your cost slide.
  2. When the web runs out of training data — how do you manufacture more? Frontier models need clean feed. When public scrape is exhausted, manufacture training-ready stock from rights-cleared document estates — same plant, different contract. We do not train on your mill uploads.
  3. Why is my local model fine but answers from my documents are still junk? Usually the model is fine and the corpus is not. A folder of files is not a knowledge base. KBMill manufactures a portable ZIP with weak material listed off the answer path—so your existing model has something fit to quote.
  4. What do “knowledge brick,” “residual-honest,” “mute,” and “stall” mean—in one sentence each? Glossary / Rosetta Stone: mill terms mapped to popular queries (portable knowledge base, hallucination controls, exclude from retrieval, local-model-docs-junk).
  5. Where do confident wrong answers from on-prem RAG come from? Wrong answers often come from unreadable scans, broken tables, and the wrong passage pulled from a messy pile. See the context Note for the failure chain.
  6. Is this another RAG product? No. RAG retrieves. Something has to make the corpus worth retrieving. KBMill prepares documents before your AI searches them. Keep your RAG; point it at a brick.
  7. Why isn’t corpus / document preparation on the on-prem LLM cost slide? Public TCO talk is loud on GPUs and quiet on making documents fit to quote. That missing line often sinks week two after the box is live.
  8. Where is the fourth leg on this on-prem ROM—host, model, RAG… and manufacture? Ask for KB/corpus manufacture as a named line. Decision makers, builders, and models often skip it. Portable question for anyone approving the spend.
  9. Why build a knowledge mill instead of another RAG layer—and why should a signer take that seriously? Incomplete stool; manufacture is the missing leg. Hard corpus, late 2023 gap, years of craft; recognition only beginning. Seriousness and proof—not pity. Leaders still own the call.
  10. How does preparing documents once save tokens, context window, and electricity? Re-cleaning messy PDFs on every question burns expensive tokens and crowds the context window. Prepare once; ask often. Fewer tokens, more room for answers.
  11. How do we get less variation so we can optimize the stack to a standard? Dump-and-hope swings failure modes. A frozen prepared package with a listed mute list and re-runnable checks—stabilizes the corpus you optimize against.
  12. What does quality mean if not a smoother demo? Answers a responsible person can stand behind after week two: mutes listed not hidden, portable leave-behind, pay only if we produce. Not trustworthy-AI slogans.
  13. How long to a corpus someone can defend—vs a multi-month data program? Speed here is time-to-defensible-corpus: drop a bounded pile, Small/Medium/Hard craft load, keep the ZIP. Hard dogfood on a commodity tower is about an hour—not a six-month priesthood.
  14. How do we build a company knowledge corpus one brick at a time—and what does initial, scale, and sustain cost? Gather domains, mill bricks, federate vs mega KB. Job prices for packages you keep; remill only what moved. Codex-quality made attainable—not only for the deepest pockets.
  15. Why mill knowledge bricks instead of one company-wide index? One melt has huge blast radius and wrong neighbors. Bricks keep human-sized scope; compose programs instead of boiling one soup.
  16. Can we keep our RAG and still use KBMill? Yes. Keep retrieval; point it at a residual-honest brick instead of the dump.
  17. We already ran a PDF extractor—why don’t we have a knowledge base? Extraction gets words out. It does not by itself produce a leave-behind with listed limits that someone will sign.
  18. Why does the model quote procedures from scanned manuals it cannot read? Image-only pages and weak OCR become ghost context. Scan craft and residual honesty matter before you trust citations.
  19. Spec tables look clean—why are the numbers invented? PDF tables often extract wrong. Models quote broken cells with confidence. Table craft is part of manufacture.
  20. What goes wrong when we dump years of SharePoint into one index? Old procedures answer new questions; wrong neighbor; unbounded blast radius. Split into bricks; don’t melt the estate.
  21. How do we know what the on-prem model is allowed to quote? If nothing is muted and nothing is listed, nobody can sign the citation. Residual-honest packages make the mute list part of the product.
  22. Is KBMill a pilot—and why are there upload caps? Yes: mutual pilot. Caps (50 files / 500 MB, one job at a time) are intentional so early problems stay bounded. We intend to extend later. Pay only if we produce.

Teaching series

Start with the plant note, then glossary and the door questions. One stall or thesis each — philosophy plus what actually runs.

  1. 0b · The plant What actually runs behind the hopper?
  2. 0c · Training stock When the web runs out — manufacture training stock
  3. 0 · Glossary Mill words and popular synonyms (Rosetta Stone).
  4. 1 · Context Where do the confident wrong answers come from?
  5. 2 · Layer RAG retrieves. Something has to make the corpus worth retrieving.
  6. 3 · The gap On-prem data work can make or break a local LLM. Why isn’t it on the cost slide?
  7. 3b · Fourth leg Where is the fourth leg on this ROM?
  8. 3c · Why a mill Why a knowledge mill — and why take it seriously?
  9. 4 · Efficiency Stop wasting tokens and power re-cleaning the same documents.
  10. 4b · Consistency Less variation. Optimize to a standard.
  11. 4c · Quality Answers you can stand behind.
  12. 4d · Speed Time to a corpus you can defend.
  13. 4e · Company corpus Build your company corpus one brick at a time.
  14. 5 · Method Why we mill knowledge bricks instead of one giant KB.
  15. 6 · Keep RAG Keep your RAG. Point it at a brick.
  16. 7 · Extract We already ran the extractor. We still don’t have a knowledge base.
  17. 8 · Scans The model is quoting procedures it cannot read.
  18. 9 · Tables Spec tables look clean. The numbers are invented.
  19. 10 · Melt We dumped ten years of SharePoint into one index.
  20. 11 · Trust How do we know what the model is allowed to quote?

Three door notes

The stall, the signature concept, and the price — then go deeper in the series above when you want the failure modes.

Why is my local model fine and the answers from my docs still junk?

Note · KBMill · the stall

You installed a local model. You dropped your manuals into chat. The answers are still wrong, empty, or confidently made up. The model is usually fine. A folder of files is not a knowledge base.

Scanned pages, website chrome, doubled letters from a bad PDF, tables that look clean and still lie — that is the pile. Splitting files so they fit a chat window is a workaround. Tools that extract text (LlamaParse and friends) stop at the extract. RAG searches whatever you stuffed in. Neither one manufactures a package you can keep and trust.

We call that package a knowledge brick: a portable ZIP you own. Known junk is muted off the answer path — taken out of what the model can quote — and listed, not hidden. That is a durable leave-behind, not a one-shot upload. You pay only if we produce. Point your existing model at the ZIP. You do not replace your stack, and you do not start a chatbot. Complexity stays in the plant. You drop files in the bin.

What does residual-honest mean?

Note · KBMill · mute, don’t hide

Some pages will not come out clean. A scan that is blank in the middle. A table whose numbers you would not bet money on. Letters doubled from a bad print layer. We could hide all of that and ship something that looks polished. We don’t.

Residual is the junk and the limits we know are still in the extract. Honest means that list is written down where you can read it. The residual board is that listed mute story — the LIBRARY_CARD plus muted chunk flags in the package — not a separate board file. Desk coverage tools stay in the plant.

A mute is a stretch we took off the answer path so the model does not treat it as fact — and so it does not invent a “typical” clause, load, or number in the hole. We call the whole stance residual-honest. That is the mill. Not an overnight rewrite by a contractor, and not “trust the PDF.” If we cannot produce an honest package, there is no charge.

Why does this cost $149, $399, or $999?

Note · KBMill · the job, not the page count

Public on-prem budgets are full of GPUs and API comparisons. They almost never put corpus fitness on the slide — the work that decides whether the local model works after the box is online. That gap is why these prices look “high” next to a parse and cheap next to a failed deployment.

The numbers look high if you think you are buying a parse. A per-page extractor will do the same PDF for a few dollars. That tool is cheaper because it is selling a different job: get the words out. We are selling a portable knowledge brick you keep. The same numbers can look “too cheap” if you just burned a GPU year and a platform project — as if a real fix must cost as much as the failure. Distrust both instincts. The price is for manufacture difficulty of an honest leave-behind, not for matching either a $3 extract or a six-figure ceremony.

Public “enterprise RAG” cost write-ups often quote tens to hundreds of thousands to build and run the retrieval system. That is a different bill from manufacturing a portable knowledge package. Our job prices can look small next to those quotes; the intent is the opposite of unserious — codex-quality corpora made attainable, domain by domain, beyond only the deepest pockets.

If you are the partner who ROMs an on-prem stand-up for a client, this is the same missing line on your estimate: hardware and “we’ll wire RAG” are already there; corpus manufacture can sit next to them as Small / Medium / Hard without inventing a new category of spend.

You get a ZIP you own — a handoffable artifact, not a session upload. On a paid mill job, open it and expect:

  • Primary Markdown — human-readable source of truth
  • Machine sidecarskb.json, chunks.jsonl, embeddings.npy (and provenance)
  • craft_brief.md — what this brick is for, and what it is not
  • Listed mutes — what we took off the answer path (LIBRARY_CARD + muted chunk flags — not hidden)
  • SECURITY_REPORT.md — what we scanned, and what we did not
  • MILL_RECEIPT.md — class and list price for your books (not a tax invoice)

We do not host your files as a library. We do not start a chatbot. Point your model at the ZIP.

Small / Medium / Hard is expected manufacture difficulty and residual craft load. A clean short manual is not the same job as a scan-heavy pile or a shelf of related books. It is not a simple page count. Take files out of the hopper and the class can move. Pay is authorized at Go and captured only when the ZIP exists. If we cannot produce a usable brick, there is no charge. After the download window we purge (72 hours).

This is not LlamaParse with a different name, not a cents-per-page parser, and not a hosted RAG chat. Those optimize for extraction speed. We optimize for a leave-behind a person or a model can actually trust. When the residual board still lists work, that is normal for real technical piles — the board is part of the product. The public mill is the factory door, not a promise of zero craft debt.

Look, don’t trust us. Public proof ZIPs on the shelf carry SECURITY_REPORT.md and craft_brief.md; muted junk stays listed on the card. Same mill as the eval above. That is what $149 / $399 / $999 is for.