Docs

Knowledge Library

How documents get into a collection, how people search them, what happens when someone asks a collection a question, and where the answers stop.

The Knowledge Library holds the documents an operation runs on: standards, work instructions, procedures, specifications, supplier paperwork, training material. People search that set, and they ask questions of it and get answers that cite the document and page the answer came from.

Collections

A collection is a set of documents that people search and ask questions of together. Every collection is scoped when it is created, and the scope decides who can reach it.

Global
Reachable across the organization. This is where the documents that do not change by location live: corporate standards, policies, procedures that are the same everywhere.
Site
Reachable at one site. This is where the paperwork that is true at that plant and nowhere else lives, so a site does not have to filter around every other site’s documents.
User
Reachable only by the person who created it. A private working set for a project, an audit, or a problem someone is chasing.

Scope is a property of the collection, not of the individual document. A document belongs to the collection it was uploaded into and inherits that collection’s reach. The Knowledge Library module page covers where this sits in the wider platform.

Adding documents

Upload by dragging files onto the upload area, or by browsing for them. One file or a batch of them.

Metadata can be set as part of the upload, and all of it is optional: a title override, which replaces the filename as the document’s display name; comma-separated tags; an author; a description.

Supported formats: PDF, .doc, .docx, .txt, .md, .csv, .xls, .xlsx, .ppt, .pptx.

The ceiling is 50MB per document. A scanned PDF is read with OCR, so a scan of a printed procedure becomes searchable text rather than a picture of a page. That matters in a plant, where the binder on the wall is often the only copy of something.

Each uploaded document is then extracted, chunked, embedded, and indexed. Search and questions reach a document once that has run.

Searching a collection

Search runs in keyword mode or semantic mode, and the mode is chosen per search.

Keyword
Exact match. Wrap words in quotes to match them as a phrase. Matching is case-insensitive. Use it when the exact term is known: a part number, a spec code, a document title.
Semantic
Concept-based, and it handles synonyms. A search for changeover time reaches a document that says setup duration. Use it when the idea is known but the wording used in the document is not.

Results can be filtered by file type, by tag, and by date range.

Every result carries a relevance score, a highlighted snippet showing the matched text in context, and its source information. Clicking a result opens the document for preview, and from there it can be downloaded or its full metadata opened.

Asking a collection a question

Chat is per collection. A question is asked of one collection, and the answer is built from the documents in it.

These limits define what a good question looks like:

  • Answers come only from the documents in the collection. Nothing outside it is consulted. If the document is not there, the answer is not there.
  • No internet access.
  • No connection to real-time systems. It does not know what a machine is doing right now, what a work order’s status is, or what a sensor read a minute ago.
  • Long conversations lose earlier context. Start a fresh conversation when the subject changes rather than carrying a long thread across topics.
  • It is not a substitute for expertise. It is a fast way to find what the documents say. Deciding what to do about that is still a person’s job.

Answers carry citations. A citation names the source document and the page or section the passage came from, and it carries a relevance score. Clicking a citation opens the exact passage, so an answer can be read against the document it came from instead of being taken on trust. That is the difference between an answer and a checkable answer.

Suggested questions are generated for a collection, so somebody opening an unfamiliar set of documents has a starting point rather than a blank box.

A conversation exports to Markdown or PDF. The export includes all citations and timestamps, which makes it something that can be attached to an investigation or handed to an auditor.

The Retrieval Engine is the module that finds the passages an answer is assembled from.

Managing documents and collections

  • Reindex a collection to rebuild its index after a batch of changes.
  • Edit document metadata at any time: title, tags, author, description.
  • Delete a document, or a whole collection.

A document cannot be moved between collections, and it cannot be edited in place. To revise one, or to relocate it: download it, edit it, upload the new version into the collection it belongs in, then delete the old one. Metadata is editable in place. The document itself is not. Plan collection scope before a large upload, because reorganizing afterwards is a download-and-re-upload job.

Analytics

Engagement
Query volume, active users, which collections get used, and peak times.
Performance
Response times, search quality, latency, and throughput.
Response quality
Thumbs ratings, how often citations are clicked, how often a question needs a follow-up, and a satisfaction measure.
Cost
Processing and storage costs, with a projected monthly figure.

The Popular Queries report is the one to read weekly. It shows what people ask most. A question asked often that keeps returning thin answers is pointing at a document nobody has uploaded, and that is the cheapest way to find the gap in what an operation has written down.

Keyboard shortcuts

Key Action
C New collection
U Upload
S Focus the search box
/ Focus the chat box
Cmd/Ctrl+K Shortcuts help
? Shortcuts
Esc Close the open panel

Scale

There is no hard cap on how many documents a collection can hold. Very large collections, meaning more than 1000 documents, may have slower search times.

If a collection is heading that way, splitting it by site or by subject keeps search fast, and it usually makes answers better too, because a narrower collection has less to be wrong about.