When the web runs out — manufacture training stock
Short answer: Frontier-class models eat clean text at industrial scale. Public web scrape is finite. When you cannot find enough good training data left to feed the next run, you do not need another scraper — you need a way to manufacture training-ready stock from document estates you actually have rights to. That is plant work.
The scarcity everyone in the room knows
Teams training or continually pre-training large models already feel the wall: high-quality public text is contested, duplicated, legally sharp, or simply used up. Throwing more raw PDFs and SharePoint dumps at the trainer does not fix it. Garbage in still makes expensive garbage — and it burns tokens and watts while doing so. See efficiency when tokens and watts bind.
The missing capability is not “one more crawl of the open web.” It is a scalable manufacturing line that turns messy, rights-cleared document piles into structured, filterable, portable training stock — with junk listed and kept off the feed when it should not be learned from.
What the plant produces for that job
The same infrastructure-class plant behind the public hopper can manufacture packages whose primary consumer is not only a RAG stack at query time — it can be a training or fine-tuning pipeline that needs clean chapters, tables recovered as data, scans made readable, and known-bad material muted and listed.
- Readable text and chunks at a human- and machine-usable grain
- Listed exclusions (mute / residual) so you do not silently train on blank scans and nav chrome
- Portable ZIP leave-behinds you can version, audit, and re-run
- Process that scales by mirroring plants, not by hoping one mega-scrape finishes
That is how you get new training content at the quality bar frontier work needs — from manuals, standards, technical libraries, and internal corpora you are allowed to use — without pretending the open web still has infinite unused gold.
Two different “training” ideas — do not mix them
This Note is easy to misread. Keep the boundary sharp:
- Your mill uploads at kbmill.com: we do not use customer job files to train our models. Isolated jobs, purge windows, pay on success — see data boundaries.
- Training stock as a product of the plant: when you (a lab, a foundation, a company with rights) need clean corpora for your training or fine-tuning runs, the plant is a way to extract and structure that stock at scale.
Same machinery. Different contract. One protects the stranger at the public door. The other serves teams whose bottleneck is “we have documents and rights — we do not have a manufacturing line to turn them into feed.”
Who this is for
Regular visitors shopping for a one-off PDF cleanup may not care. People trying to feed frontier-class training runs — or to stand up continual training on domain libraries — will care if there is a serious, copyable process instead of a pile of brittle extract scripts.
Related legitimacy: what runs behind the hopper, packages not one melt, why a mill.
What to do
If you are evaluating training-stock manufacture: open a sample package, judge whether the leave-behind is feed-grade for your pipeline, then talk to the operator about bounded pilots or program-scale plant copies — not only the public Small/Medium/Hard door.