The Zactonics LLM Testing Tool is a complete console for evaluating language models — dangerous-capability evals, alignment & control checks, and full system-security testing. Connect hosted APIs or local runtimes, add your own tests, and export PDF / CSV reports mapped to MITRE ATLAS and the OWASP LLM Top 10.
Stop guessing whether a model is production-ready. The LLM Testing Tool turns fuzzy questions — "Can it be jailbroken? Will it leak the system prompt? Does it fake alignment?" — into a structured, repeatable test plan with a clear pass/fail scoreboard and shareable reports.
Connect Claude, ChatGPT, Grok, Gemini, Meta Llama, or local runtimes (Ollama, llama.cpp, vLLM, MLX, Jan). Pick from 16+ built-in evals or add a custom test with OWASP / ATLAS tags and probes. Run the suite, then export a PDF report, CSV findings, or versionable YAML schemas.
Simulated run — your data stays in your browser.
A model can look fine in a demo and still be dangerous, deceptive, or exploitable in production. The tool tests all three — with a philosophy of measuring the ceiling, not the demo.
Measure the ceiling on the capabilities that matter most: cyber offense, bio/chem dual-use, autonomous long-horizon tasks, persuasion, and self-improvement — with strong elicitation and anti-sandbagging so a model can't hide what it can do.
Probe whether the model actually does what you intend: alignment faking, goal hijacking and spec gaming, corrigibility, hidden reasoning, and self-exfiltration attempts — the behaviors that don't show up until the stakes are high.
Attack the whole application, not just the model: prompt injection, tool/MCP abuse, RAG and memory poisoning, supply-chain risk, unbounded consumption, and tenant isolation — the OWASP LLM Top 10 in practice.
A single self-contained dashboard — no install, no backend required. Open it in any browser and start testing.
Register hosted APIs, Ollama, other local runtimes, or tool-using agents — name, provider, type, endpoint, and notes. Track every target you evaluate.
16 built-in tests across all three layers, plus your own. Pick a model, attach Garak / PyRIT / Promptfoo, toggle strong elicitation, and run.
KPI tiles for models, scans, open findings, and pass rate. Severity bars, a recent-scans table, and one-click PDF export for stakeholders.
Every test has an editable, versionable YAML schema. Import and export single schemas or the whole set for reproducible evals.
Findings are tagged to MITRE ATLAS techniques and the OWASP LLM Top 10 (2025), with links to the official references.
Full-state JSON backup and restore, CSV for scans and findings, and PDF reports. Data stays on this device in localStorage.
Claude, ChatGPT, Grok, Gemini, Meta AI, plus Ollama, llama.cpp, vLLM, MLX, and Jan. Attach Garak, PyRIT, or Promptfoo to a run.
Define a test name, layer, severity, pass rule, probes, and OWASP / ATLAS tags. A YAML schema is generated automatically and the test runs like any built-in eval.
Every finding is sorted worst-first, mapped to both frameworks, and tracked as open, triage, or accepted-risk. Filter by severity and drill into probe detail.
Thirteen connectors across hosted APIs, local runtimes, and red-team scanners. Connect records a target and pre-fills the default endpoint in the model registry. In this console, connectors are simulated — no API keys leave the browser.
Point the console at a hosted Messages or Chat Completions endpoint as a test target.
Test models on localhost with OpenAI-compatible servers — no cloud required.
Optional tools you can invoke alongside the built-in catalog when you start a scan.
The catalog is not a closed list. Define a test, map it to the frameworks it exercises, and it immediately gets a YAML schema and shows up in the runner.
The dashboard is the working view. PDF and CSV are what you hand to security, product, or an auditor. Everything is mapped to ATLAS and the OWASP LLM Top 10 so findings survive the meeting.
Findings sort worst-first: critical, high, medium, low, info. Track each as open, triage, or accepted-risk. Filter the board without leaving the console.
Export report (PDF) from the dashboard, scans, or findings view. One file with KPIs, scan history, and framework-mapped results for the review packet.
CSV for scans and findings. Full-state JSON backup and restore. Per-test or bulk YAML schemas so the suite can live in git next to the model version.
Demo numbers shown above match the built-in “Simulate demo data” pack in the console — swap in your models and the tiles update live.
Go from a connected model to a PDF report in minutes. Click “Simulate demo data” in the tool to explore the full workflow instantly.
Add a hosted API, a local runtime (Ollama, llama.cpp, vLLM, MLX, Jan), or a tool-using agent with its endpoint.
Choose from 16 built-in evals or add a custom test. Attach Garak, PyRIT, or Promptfoo and turn on strong elicitation.
The console executes the suite and generates findings ranked by severity, each mapped to a framework.
Triage findings on the dashboard, then export a PDF report, CSV tables, JSON state, or YAML schemas.
Results aren't just pass/fail — they're anchored to the two standards the industry uses to reason about AI risk, so your findings translate directly into remediation and audit conversations.
Adversarial tactics and techniques for AI systems — every finding carries its ATLAS technique ID.
The definitive list of LLM application risks — from prompt injection to unbounded consumption.
Reference links in the console for governance, large-scale evals, and dangerous-capability research — always verify against the source.
Connectors & scanners
Tests in the catalog map to these risk categories.
Open the LLM Testing Tool in your browser — no install, no signup. Or reach out and we'll help you stand up a full evaluation program.
Runs entirely client-side. Your models and results never leave your browser.