The idea
Extractium gathers what an organization already publishes into one compendium: a searchable collection of its content, written as static files. Point it at a website, a knowledge base portal, GitHub repositories, a YouTube channel, a library repository, or a folder of files. It gathers the content, prepares it for both keyword and meaning-based search, and writes it out in several formats.
How it works
Unlike a vector database, Extractium needs no server, no database, and no API to run. Every output is a static file that can be hosted anywhere, including GitHub Pages, and the same build feeds all of them at once: a search index, llms.txt files for AI assistants, a SQLite database, and a folder of Markdown. Sources and outputs are plug-ins.
Context
Extractium grew out of the indexing engine in Field Station AI, which remains an example front end for the JSON output. I wrote it as part of my work at the Depression Center. The project is university property and open source.
