FAQ

Answers to common questions about Cortyxia.

Get clear answers about how Cortyxia handles AI memory, context retrieval, privacy, self-hosting, and cost reduction. Whether you are evaluating the platform for a team or deploying it across your organization, these questions cover the details that matter most.

Common Questions

Straight answers on how Cortyxia saves you money, keeps your data yours, and makes production AI actually work.

Cortyxia sits between your application and your model provider as a high-performance proxy. Point your base URL at Cortyxia instead of OpenAI, Anthropic, or whoever you use today, and every LLM call flows through us automatically. We retrieve the right memory, inject it into the prompt, trim what does not belong, and forward the optimized request to your provider. No rebuild. No new workflow. Your team keeps building exactly how they already do, but every call gets smarter context and full visibility from day one.
Most AI tools forget everything the moment a session ends. That means your team re-explains the same context over and over, and every department runs on a different version of the truth. Cortyxia turns scattered interactions into one living memory graph. Pool keys when you want cross-team intelligence, like a decision made in Cursor surfacing in Salesforce the same day. Keep keys isolated when a project needs strict boundaries. You choose what converges and what stays private, so knowledge compounds instead of decaying.
Yes. For teams with data residency, compliance, or security requirements, Cortyxia deploys fully on-premise or inside your VPC. The core proxy runs on SQLite by default. PostgreSQL handles telemetry and analytics. Redis is optional for caching. Your memory, your queries, and your audit trail stay inside your perimeter. You get enterprise-grade AI memory without handing custody to a third-party cloud.
Retrieval adds some overhead before the provider call, because Cortyxia assembles memory at inference instead of blindly replaying history. In return, prompts stay bounded and you send far fewer tokens. On our published evals, that trade paid off in spend and task success. Exact latency depends on workload, retrieval temperature, and provider.
Yes, when you were paying for full-context replay or bloated prompts. On our published 50-question enterprise governance eval, Cortyxia cut prompt tokens by 80.8% versus full-context with quality held, compounding to 10.2× fewer tokens by question 50. A 20-turn IDE session saw 91.5% token reduction with comparable code quality. SWE-style fixes used 70% fewer tokens while resolving 100% of tasks versus 73.3% for full-context. You pay for what the query needs, not the entire history.
Cortyxia is model-agnostic by design. OpenAI, Anthropic, Google Gemini, DeepSeek, xAI, Groq, and more route through the same memory layer. Switch providers or run different models for different tasks without rebuilding your context pipeline. Your memory stays put. Your provider becomes a choice, not a lock-in.
Everything you need to run AI in production with confidence, in one pane. Compare models on cost, latency, and six quality metrics. Score every prompt and reply. Auto-detect guardrails from your instructions and get alerted when something breaks. Trace tool calls, memory lookups, and agent steps on every message. No stitching Datadog, prompt labs, and compliance spreadsheets together. You see what happened, what it cost, and what went wrong before your users do.
Minutes, not a quarter-long integration project. Create a project, generate an API key, and point your application's base URL at Cortyxia. SDK, CLI, and coding agents across OpenAI, Anthropic, Gemini, and other providers work out of the box. Memory starts capturing on the first call. Most teams are running in a dev environment the same day and pushing to production once they see the token savings and retrieval quality for themselves.