Deterministic runtime for agent evaluation
-
Updated
Mar 25, 2026 - Python
Deterministic runtime for agent evaluation
KayaDB distributed key-value storage engine
Zero-LLM deterministic jailbreak regression benchmark: runs a repeatable battery of known-pattern attacks against an LLM endpoint and scores refusal/partial/compliance across runs. Catch when a model or prompt update silently got weaker. Single-turn, responsible-use. pip install hermes-jailbench
Local-first C# project for deterministic prompt versioning, A/B evaluation, and evidence-based promotion using structured scoring.
Deterministic UI testing framework for TS/JS — canvas-first. Virtual time, scene contracts, model-driven discovery.
Deterministic simulation testing framework for distributed systems in .NET
WordleBench — Deterministic AI Wordle benchmark. Compare 34+ LLMs (GPT-5, Claude 4.5, Gemini, Grok, Llama) head-to-head on accuracy, speed, and cost across 50 standardized words.
A general harness for repeatable app states.
LogOS A closed-loop cognitive Operating surface: written in Rust, deployed via self-verifying 'narrow-waist' Nix OS + Mirage OS Uniquernel to Google Cloud Run/Kubernetes,
Reproduce OpenAI Agents SDK runtime bugs without API keys — real Runner, scripted models, verifiable certificates.
A Raft consensus library built from scratch in Rust, with a deterministic network simulator for testing leader election, log replication, and partition recovery.
The deterministic heap groomer for C/C++ memory debugging.
Deterministic Rust testing utility for simulation and stochastic workflows
Type-safe clock abstractions for Go with zero dependencies
A small public exemplar of governed AI-assisted delivery: bounded work, explicit authority, deterministic verification, evidence, and human decision.
Evidence-first Excel financial-model review with deterministic checks, finance tie-outs, cell-level proof, cited memos, and Markdown export.
Deterministic simulation testing skill: fault injection, shrinkable replays, and self-verifying docs.
A black-box evaluator that checks agent behavior and evidence independently, then produces reproducible reports.
Python CLI for environmental marketing-claim risk review with structured outputs, image support, and offline mock testing.
Add a description, image, and links to the deterministic-testing topic page so that developers can more easily learn about it.
To associate your repository with the deterministic-testing topic, visit your repo's landing page and select "manage topics."