Skip to content

feat(appkit): agent eval framework, judge, mlflow connector (stack 2/5) - #478

Open
MarioCadenas wants to merge 5 commits into
mainfrom
pr/agent-evals-2-framework
Open

feat(appkit): agent eval framework, judge, mlflow connector (stack 2/5)#478
MarioCadenas wants to merge 5 commits into
mainfrom
pr/agent-evals-2-framework

Conversation

@MarioCadenas

Copy link
Copy Markdown
Collaborator

Stack 2/5 · targets pr/agent-evals-1-tracing (review after #1).

The core eval framework, plus LLM-as-judge and the MLflow REST connector.

  • Authoring (defineEval): drive an agent over HTTP against a running app; assert with t.succeeded(), t.calledTool(), t.check(value, matcher) (includes/equals/matches). Gate-by-default, .soft() to demote.
  • Native MLflow Evaluation runs: a run tagged genai_evaluate; each turn's trace links via mlflow.sourceRun; per-assertion feedback written via the assessments REST API.
  • LLM-as-judge (t.judge.factuality/closedQA/custom) via autoevals → a Databricks serving endpoint.
  • connectors/mlflow: MlflowClient (host/token, post/postResult, serving URL) + resolveDatabricksAuth/resolveWorkspaceClient (OAuth from a CLI profile — no hand-set PAT). Extracted so both evals and future callers share the REST/auth layer.
  • appkit agent eval CLI.

Squashed history note: contains the framework, judge, and connector-extraction commits.

@MarioCadenas
MarioCadenas requested a review from a team as a code owner July 16, 2026 14:26
@MarioCadenas
MarioCadenas requested review from pkosiec and removed request for a team July 16, 2026 14:26
@MarioCadenas
MarioCadenas force-pushed the pr/agent-evals-1-tracing branch from 412b08d to f517306 Compare August 5, 2026 15:22
Base automatically changed from pr/agent-evals-1-tracing to main August 13, 2026 09:27
…runs

eve-style eval authoring (defineEval + t-context + matchers) discovered from
config/agents/<id>/evals/*.eval.ts and run via 'appkit agent eval' against a
running app. Streams per-eval progress and gates CI via exit code.

When Databricks creds + an experiment are set, it creates a real MLflow
evaluation run (mlflow.runType=genai_evaluate): each eval's trace links to the
run, pass/fail is written as feedback assessments, and aggregate metrics are
logged. All via the MLflow REST API.

Signed-off-by: MarioCadenas <MarioCadenas@users.noreply.github.com>
Extend the agent eval framework and tighten MLflow output to match the
native `mlflow.genai.evaluate` experience:

- LLM-as-judge via autoevals (factuality, closedQA, custom), pointed at a
  Databricks serving endpoint; exposed through `t.judge.*`.
- One Feedback assessment per assertion (judges as LLM_JUDGE with score +
  rationale) plus an overall `appkit_eval`; assessment names sanitized to
  `[A-Za-z0-9_-]` since the API rejects dots.
- Trace-table parity: set Request/Response previews and the `mlflow.traceName`
  tag (the Trace-name column reads the tag, not the span name).
- Eval runs carry `mlflow.source.name`/`type` tags so linked traces show
  Source and Run name; live chat traces have no run so those stay empty.
- Example judge eval under config/agents/query/evals.

Signed-off-by: MarioCadenas <MarioCadenas@users.noreply.github.com>
Introduce connectors/mlflow as the shared REST + auth layer for MLflow,
so the eval runner (and future callers) stop threading host/token and
hand-rolling fetch/URL logic:

- MlflowClient owns {host, token}: normalizes the host once, exposes
  post() (throws) for runs/* and postResult() (structured failure) for
  best-effort assessment writes, plus servingEndpointsUrl() for the judge.
- resolveDatabricksAuth() mints an OAuth bearer from a CLI profile via the
  SDK WorkspaceClient (the AppKit-native path), so `agent eval` no longer
  requires a hand-set DATABRICKS_TOKEN. Adds an `--profile` flag.
- Eval run create/finish, assessment reporting, and the judge take the
  client; the agents plugin's host normalization now delegates to the
  connector's normalizeHost.

The mlflow-tracing SDK wrapper stays in the agents plugin: it manages a
process-global provider (like TelemetryManager) and has an agent-shaped
API, so it isn't a connector.

Signed-off-by: MarioCadenas <MarioCadenas@users.noreply.github.com>
Signed-off-by: MarioCadenas <MarioCadenas@users.noreply.github.com>
Post-rebase integration with main's biome->oxc migration (#538) and the
SDK-facade boundary rule (#534):
- Route the mlflow connector's auth through createWorkspaceClient instead
  of importing @databricks/sdk-experimental directly (oxlint
  no-restricted-imports); behaviour is unchanged.
- Apply oxfmt import grouping to the evals + connector files authored
  before the migration.

Signed-off-by: MarioCadenas <MarioCadenas@users.noreply.github.com>
@MarioCadenas
MarioCadenas force-pushed the pr/agent-evals-2-framework branch from 0b55719 to 1431ce7 Compare August 18, 2026 12:57
@github-actions

Copy link
Copy Markdown
Contributor

📦 Bundle size report

Compared against bundle-size-baseline.json (main).

@databricks/appkit

npm tarball (packed): 877 KB (+37 KB) — gzipped download (dist + bin; excludes release-only docs/NOTICE).

dist raw gzip
JS (runtime) 899 KB (+30 KB) 316 KB (+13 KB)
Type declarations 335 KB (+20 KB) 118 KB (+9.2 KB)
Source maps 1.7 MB (+60 KB) 591 KB (+25 KB)
Other 11 KB 3.7 KB
Total 3.0 MB (+110 KB) 1.0 MB (+47 KB)
Per-entry composition (own code — deps external (as shipped))
Entry Initial (gz) Lazy (gz) Total (gz) node_modules (min) Own code (min)
. 88 KB (+3 B) 2.5 KB 91 KB (+3 B) external 288 KB
./beta 53 KB (+4.1 KB) 457 B 53 KB (+4.1 KB) external 154 KB (+11 KB)
./type-generator 21 KB 0 B 21 KB external 61 KB

Chunks:

Entry Chunk Load Size (gz)
. index.js initial 84 KB
. utils.js initial 4.0 KB
. remote-tunnel-manager.js lazy 2.5 KB
./beta beta.js initial 37 KB
./beta stream-manager.js initial 5.8 KB
./beta wide-event-emitter.js initial 3.2 KB
./beta databricks.js initial 3.0 KB
./beta configuration.js initial 2.1 KB
./beta service-context.js initial 1.3 KB
./beta client.js initial 434 B
./beta client-options.js initial 220 B
./beta supervisor-api.js lazy 192 B
./beta databricks.js lazy 142 B
./beta index.js lazy 123 B
./type-generator index.js initial 21 KB

@databricks/appkit-ui

npm tarball (packed): 342 KB (-291 B) — gzipped download (dist + bin; excludes release-only docs/NOTICE).

dist raw gzip
JS (runtime) 390 KB 130 KB (+1 B)
Type declarations 228 KB 83 KB
Source maps 752 KB (-334 B) 247 KB (-197 B)
CSS 16 KB (-462 B) 3.2 KB (-90 B)
Total 1.4 MB (-796 B) 464 KB (-286 B)
Per-entry composition (consumer bundle — deps bundled, peerDeps external)
Entry Initial (gz) Lazy (gz) Total (gz) node_modules (min) Own code (min)
./js 5.3 KB 49 KB 55 KB 208 KB 14 KB
./js/beta 20 B 0 B 20 B 0 B 0 B
./react 432 KB (+127 B) 49 KB 480 KB (+127 B) 1.3 MB 175 KB
./react/beta 1.0 KB 0 B 1.0 KB 0 B 1.9 KB

Chunks:

Entry Chunk Load Size (gz)
./js index.js initial 5.2 KB
./js chunk initial 120 B
./js apache-arrow lazy 49 KB
./js/beta beta.js initial 20 B
./react index.js initial 430 KB
./react tslib initial 2.1 KB
./react apache-arrow lazy 49 KB
./react/beta beta.js initial 1.0 KB

@github-actions

Copy link
Copy Markdown
Contributor

🤖 AppKit PR bot

🔬 Run evals

Start an eval for this PR from the evals-monitor app: Go to Evals Monitor →

📦 Try this PR's app template

Scaffolds a new app from this PR's SDK build. Run it in any folder (requires the GitHub CLI — gh auth login — and the Databricks CLI):

gh run download 32139644522 -R databricks/appkit -n appkit-template-0.61.1-pr.0af9785-pr-agent-evals-2-framework-478 -D appkit-pr-478 \
  && unzip -o "appkit-pr-478/appkit-template-0.61.1-pr.0af9785-pr-agent-evals-2-framework-478.zip" -d "appkit-pr-478" \
  && databricks apps init --template "appkit-pr-478"

The template pins @databricks/appkit and @databricks/appkit-ui to tarballs built from this branch, so the scaffolded app runs against this PR's code.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant