Appearance
Attunement System Architecture & Engineering
Technical specification of the underlying infrastructure, services, and processing pipelines supporting ToFF Attunement.
1. Container Topology
ToFF Attunement operates as a sovereign multi-container system orchestrating data ingestion, search indexing, media processing, and visual rasterization:
| Container | Image / Technology | Purpose | Port / Scope |
|---|---|---|---|
karakeep-web | Node.js / Next.js Core | Web UI, API gateway, and business logic | 3000 (Internal) |
karakeep-meili | Meilisearch v1.41 | Sub-second full-text and vector search engine | 7700 (Internal) |
karakeep-chrome | Chromium Headless Shell | Full-page visual rasterization and PDF generation | 9222 (Internal CDP) |
caddy | Caddy v2 Reverse Proxy | Automated Let's Encrypt TLS termination and edge routing | 80, 443 (Public) |
2. Autonomous Video Ingress & Preservation Engine
Permanent Preservation Guarantee
ToFF Attunement incorporates an in-house media extraction engine powered by an integrated background yt-dlp daemon and internal object storage. It automatically captures full-resolution offline copies of Instagram Reels, YouTube, TikTok, and X videos.
Media Processing Lifecycle:
- Link Ingestion: The system detects an external video URL submitted via browser extension, mobile share sheet, or web portal.
- Extraction Worker: The background daemon invokes
yt-dlp, bypasses transient platform restrictions, and pulls the maximum resolution stream (bestvideo+bestaudio/best). - Local Storage Commit: The resulting video container is committed directly to Foundation-controlled encrypted volumes.
- Thumbnail & Metadata Indexing: High-definition poster frames and creator captions are extracted and indexed for full-text search.
- Streaming Player Delivery: The asset is immediately made available for in-browser HTML5 playback without requiring full local downloads.
3. Visual Intelligence & AI Tagging Architecture
When a webpage or screenshot is ingested:
- Rasterization: Chromium Headless connects via Chrome DevTools Protocol (CDP) to render the target URL in a simulated 1280x800 desktop or mobile viewport.
- Optical Character Recognition (OCR): Embedded Tesseract engines scan rasterized images to extract typography, signage, and embedded text.
- Semantic Tagging: OpenAI vision and text models synthesize descriptive metadata, categorize themes, and assign taxonomy tags (
#Instagram,#Press, etc.).
4. Search & Indexing Engine
The platform utilizes Meilisearch to maintain inverted indices across:
- Title and description fields
- Extracted OCR text from screenshots and photos
- Full webpage readability bodies
- Curatorial notes and user tags
- Chronological timestamps and domain namespaces