feat(appkit): agent eval framework, judge, mlflow connector (stack 2/5) - #478
Open
MarioCadenas wants to merge 5 commits into
Open
feat(appkit): agent eval framework, judge, mlflow connector (stack 2/5)#478MarioCadenas wants to merge 5 commits into
MarioCadenas wants to merge 5 commits into
Conversation
MarioCadenas
force-pushed
the
pr/agent-evals-1-tracing
branch
from
August 5, 2026 15:22
412b08d to
f517306
Compare
…runs eve-style eval authoring (defineEval + t-context + matchers) discovered from config/agents/<id>/evals/*.eval.ts and run via 'appkit agent eval' against a running app. Streams per-eval progress and gates CI via exit code. When Databricks creds + an experiment are set, it creates a real MLflow evaluation run (mlflow.runType=genai_evaluate): each eval's trace links to the run, pass/fail is written as feedback assessments, and aggregate metrics are logged. All via the MLflow REST API. Signed-off-by: MarioCadenas <MarioCadenas@users.noreply.github.com>
Extend the agent eval framework and tighten MLflow output to match the native `mlflow.genai.evaluate` experience: - LLM-as-judge via autoevals (factuality, closedQA, custom), pointed at a Databricks serving endpoint; exposed through `t.judge.*`. - One Feedback assessment per assertion (judges as LLM_JUDGE with score + rationale) plus an overall `appkit_eval`; assessment names sanitized to `[A-Za-z0-9_-]` since the API rejects dots. - Trace-table parity: set Request/Response previews and the `mlflow.traceName` tag (the Trace-name column reads the tag, not the span name). - Eval runs carry `mlflow.source.name`/`type` tags so linked traces show Source and Run name; live chat traces have no run so those stay empty. - Example judge eval under config/agents/query/evals. Signed-off-by: MarioCadenas <MarioCadenas@users.noreply.github.com>
Introduce connectors/mlflow as the shared REST + auth layer for MLflow,
so the eval runner (and future callers) stop threading host/token and
hand-rolling fetch/URL logic:
- MlflowClient owns {host, token}: normalizes the host once, exposes
post() (throws) for runs/* and postResult() (structured failure) for
best-effort assessment writes, plus servingEndpointsUrl() for the judge.
- resolveDatabricksAuth() mints an OAuth bearer from a CLI profile via the
SDK WorkspaceClient (the AppKit-native path), so `agent eval` no longer
requires a hand-set DATABRICKS_TOKEN. Adds an `--profile` flag.
- Eval run create/finish, assessment reporting, and the judge take the
client; the agents plugin's host normalization now delegates to the
connector's normalizeHost.
The mlflow-tracing SDK wrapper stays in the agents plugin: it manages a
process-global provider (like TelemetryManager) and has an agent-shaped
API, so it isn't a connector.
Signed-off-by: MarioCadenas <MarioCadenas@users.noreply.github.com>
Signed-off-by: MarioCadenas <MarioCadenas@users.noreply.github.com>
Post-rebase integration with main's biome->oxc migration (#538) and the SDK-facade boundary rule (#534): - Route the mlflow connector's auth through createWorkspaceClient instead of importing @databricks/sdk-experimental directly (oxlint no-restricted-imports); behaviour is unchanged. - Apply oxfmt import grouping to the evals + connector files authored before the migration. Signed-off-by: MarioCadenas <MarioCadenas@users.noreply.github.com>
MarioCadenas
force-pushed
the
pr/agent-evals-2-framework
branch
from
August 18, 2026 12:57
0b55719 to
1431ce7
Compare
Contributor
📦 Bundle size reportCompared against
|
| dist | raw | gzip |
|---|---|---|
| JS (runtime) | 899 KB (+30 KB) | 316 KB (+13 KB) |
| Type declarations | 335 KB (+20 KB) | 118 KB (+9.2 KB) |
| Source maps | 1.7 MB (+60 KB) | 591 KB (+25 KB) |
| Other | 11 KB | 3.7 KB |
| Total | 3.0 MB (+110 KB) | 1.0 MB (+47 KB) |
Per-entry composition (own code — deps external (as shipped))
| Entry | Initial (gz) | Lazy (gz) | Total (gz) | node_modules (min) | Own code (min) |
|---|---|---|---|---|---|
. |
88 KB (+3 B) | 2.5 KB | 91 KB (+3 B) | external | 288 KB |
./beta |
53 KB (+4.1 KB) | 457 B | 53 KB (+4.1 KB) | external | 154 KB (+11 KB) |
./type-generator |
21 KB | 0 B | 21 KB | external | 61 KB |
Chunks:
| Entry | Chunk | Load | Size (gz) |
|---|---|---|---|
. |
index.js |
initial | 84 KB |
. |
utils.js |
initial | 4.0 KB |
. |
remote-tunnel-manager.js |
lazy | 2.5 KB |
./beta |
beta.js |
initial | 37 KB |
./beta |
stream-manager.js |
initial | 5.8 KB |
./beta |
wide-event-emitter.js |
initial | 3.2 KB |
./beta |
databricks.js |
initial | 3.0 KB |
./beta |
configuration.js |
initial | 2.1 KB |
./beta |
service-context.js |
initial | 1.3 KB |
./beta |
client.js |
initial | 434 B |
./beta |
client-options.js |
initial | 220 B |
./beta |
supervisor-api.js |
lazy | 192 B |
./beta |
databricks.js |
lazy | 142 B |
./beta |
index.js |
lazy | 123 B |
./type-generator |
index.js |
initial | 21 KB |
@databricks/appkit-ui
npm tarball (packed): 342 KB (-291 B) — gzipped download (dist + bin; excludes release-only docs/NOTICE).
| dist | raw | gzip |
|---|---|---|
| JS (runtime) | 390 KB | 130 KB (+1 B) |
| Type declarations | 228 KB | 83 KB |
| Source maps | 752 KB (-334 B) | 247 KB (-197 B) |
| CSS | 16 KB (-462 B) | 3.2 KB (-90 B) |
| Total | 1.4 MB (-796 B) | 464 KB (-286 B) |
Per-entry composition (consumer bundle — deps bundled, peerDeps external)
| Entry | Initial (gz) | Lazy (gz) | Total (gz) | node_modules (min) | Own code (min) |
|---|---|---|---|---|---|
./js |
5.3 KB | 49 KB | 55 KB | 208 KB | 14 KB |
./js/beta |
20 B | 0 B | 20 B | 0 B | 0 B |
./react |
432 KB (+127 B) | 49 KB | 480 KB (+127 B) | 1.3 MB | 175 KB |
./react/beta |
1.0 KB | 0 B | 1.0 KB | 0 B | 1.9 KB |
Chunks:
| Entry | Chunk | Load | Size (gz) |
|---|---|---|---|
./js |
index.js |
initial | 5.2 KB |
./js |
chunk |
initial | 120 B |
./js |
apache-arrow |
lazy | 49 KB |
./js/beta |
beta.js |
initial | 20 B |
./react |
index.js |
initial | 430 KB |
./react |
tslib |
initial | 2.1 KB |
./react |
apache-arrow |
lazy | 49 KB |
./react/beta |
beta.js |
initial | 1.0 KB |
Contributor
🤖 AppKit PR bot🔬 Run evalsStart an eval for this PR from the evals-monitor app: Go to Evals Monitor → 📦 Try this PR's app templateScaffolds a new app from this PR's SDK build. Run it in any folder (requires the GitHub CLI — gh run download 32139644522 -R databricks/appkit -n appkit-template-0.61.1-pr.0af9785-pr-agent-evals-2-framework-478 -D appkit-pr-478 \
&& unzip -o "appkit-pr-478/appkit-template-0.61.1-pr.0af9785-pr-agent-evals-2-framework-478.zip" -d "appkit-pr-478" \
&& databricks apps init --template "appkit-pr-478"The template pins |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stack 2/5 · targets
pr/agent-evals-1-tracing(review after #1).The core eval framework, plus LLM-as-judge and the MLflow REST connector.
defineEval): drive an agent over HTTP against a running app; assert witht.succeeded(),t.calledTool(),t.check(value, matcher)(includes/equals/matches). Gate-by-default,.soft()to demote.genai_evaluate; each turn's trace links viamlflow.sourceRun; per-assertion feedback written via the assessments REST API.t.judge.factuality/closedQA/custom) via autoevals → a Databricks serving endpoint.connectors/mlflow:MlflowClient(host/token, post/postResult, serving URL) +resolveDatabricksAuth/resolveWorkspaceClient(OAuth from a CLI profile — no hand-set PAT). Extracted so both evals and future callers share the REST/auth layer.appkit agent evalCLI.Squashed history note: contains the framework, judge, and connector-extraction commits.