Orlok

Importing documents

File office documents and scans into the wiki.

Office managers and administrators upload office documents (PDF, Word, Excel, PowerPoint, OpenDocument, HTML, CSV, text) from the wiki import page and choose one of their own colleagues in that office to file them. Each step shows while it happens: reading the text, reading the images, filing. The colleague first records what the document brings (new knowledge, an update, a contradiction, or nothing, in which case it is not filed), then writes a source page and updates or creates the topic pages it touches, linked together.

  • Scanned pages and images are read on the server with OCR (Tesseract, Italian and English, no network needed), and the import says how many pages needed it. Unreadable words are marked [illegible], never guessed. An administrator can also choose a model that reads images, for better quality: those images are then sent to its provider.
  • The extracted text is filed by the colleague with the model of its profile, so it goes to that model's provider, unless it is a local model with Ollama.
  • The same file is imported only once per office: a second upload shows as already imported.
  • The original file and the extracted text are kept, so a document can be filed again with File again, for example after choosing a model that reads images.
  • Removing a document can also delete the pages written only from it: shared pages just lose it as a source, and links to deleted pages are removed.