Appearance
Visual Intelligence & AI Tagging Architecture
Technical Specification ARCH-AI-01
Managed by CultureOS
ToFF Attunement incorporates an automated multi-stage processing pipeline to transform raw web links and social posts into structured, semantically searchable institutional intelligence.
The pipeline runs entirely in isolated background worker containers on Foundation server infrastructure, configured, deployed, and managed by CultureOS.
Processing Pipeline Overview
When an asset is ingested via browser extension, mobile application, or API, it passes through four sequential phases:
[ Ingress Event ]
|
v
[ Headless Chrome Crawler ] --> Generates Full-Page Screenshots & Clean Readability DOM
|
v
[ Visual OCR & Media Engine ] --> Extracts Text from Images & Embedded Graphics
|
v
[ Multimodal AI Inference ] --> Performs Automated Tagging & Editorial Summary
|
v
[ MeiliSearch Indexer ] --> Generates Vector Embeddings for Hybrid Semantic RetrievalTechnical Subsystems
1. Headless Browser Crawling
- Container:
karakeep-chrome - Mechanism: Dedicated Chromium instance controlled via Chrome DevTools Protocol (CDP).
- Output:
- Clean article Markdown text stripped of advertisements and navigation clutter.
- Full-resolution viewport PNG snapshot for permanent visual record keeping.
- OpenGraph metadata (author, published timestamp, lead image).
2. Optical Character Recognition (OCR)
- Automated text extraction scans all captured images, photographs of gallery placards, and flyer graphics.
- Extracted textual strings are injected into the global full-text search index, allowing researchers to search for words embedded inside images even if no written caption exists.
3. Automated Taxonomy & Semantic Tagging
The platform can interface with OpenAI-compatible inference endpoints (cloud-hosted or self-hosted via Ollama) to assign metadata tags automatically:
| Dimension | Processing Function |
|---|---|
| Automatic Tagging | Evaluates article semantics against institutional guidelines and generates topical tags. |
| Abstractive Summary | Produces a 2-3 sentence executive synopsis stored alongside the primary record. |
| Contextual Embeddings | Generates high-dimensional vector representations enabling conceptual similarity search. |
Model Integration & Configuration Guide
ToFF Attunement provides a native inference abstraction layer supporting three primary deployment architectures:
Option 1: Direct OpenAI Cloud Gateway (Recommended)
This configuration delivers maximum semantic fidelity, rapid turnaround, and robust multimodal vision parsing at minimal operational cost:
- Access the Foundation server environment:bash
ssh toff-oracle - Append the following inference directives to
/opt/karakeep/.env:bash# Primary Model Gateway OPENAI_API_KEY=sk-proj-... # Inference Model Specifications INFERENCE_TEXT_MODEL=gpt-4o-mini INFERENCE_IMAGE_MODEL=gpt-4o-mini # Behavioral Flags INFERENCE_ENABLE_AUTO_TAGGING=true INFERENCE_ENABLE_AUTO_SUMMARIZATION=true CHAT_ENABLED=true SEMANTIC_SEARCH_ENABLED=true - Restart the Attunement service cluster:bash
cd /opt/karakeep && docker compose up -d --force-recreate karakeep-web
Option 2: OpenAI-Compatible Third-Party Endpoints (DeepSeek / OpenRouter / Groq)
If the Foundation elects to utilize specialized models or high-throughput external providers:
bash
OPENAI_BASE_URL=https://api.deepseek.com/v1
OPENAI_API_KEY=sk-...
INFERENCE_TEXT_MODEL=deepseek-chat
INFERENCE_IMAGE_MODEL=gpt-4o-mini
INFERENCE_ENABLE_AUTO_TAGGING=true
INFERENCE_ENABLE_AUTO_SUMMARIZATION=true
CHAT_ENABLED=trueOption 3: Sovereign Air-Gapped Local Inference (Ollama)
For completely air-gapped environments requiring zero external API transmission:
bash
OLLAMA_BASE_URL=http://localhost:11434
INFERENCE_TEXT_MODEL=llama3.2
INFERENCE_IMAGE_MODEL=llava
INFERENCE_ENABLE_AUTO_TAGGING=true
INFERENCE_ENABLE_AUTO_SUMMARIZATION=trueData Privacy & Telemetry Guard
To preserve Foundation confidentiality:
- No donor records, financial instruments, or internal legal agreements from ToFF Sign ever enter this processing pipeline.
- Ingress into ToFF Attunement is restricted exclusively to public domain materials, cultural articles, and open community discourse.