
Evals Create Suite
- 2 installs
- 21.2k repo stars
- Updated August 5, 2026
- elastic/kibana
evals-create-suite skill documents Scaffold a new LLM evaluation suite package with Playwright config, evaluate fixture, and package files.
About
evals-create-suite skill documents Scaffold a new LLM evaluation suite package with Playwright config, evaluate fixture, and package files. Use when creating a new eval suite, adding an evals package for a plugin, or setting up the boilerplate for offline LLM evaluations.. name: evals-create-suite disable-model-invocation: true
- Scaffold a new LLM evaluation suite package with Playwright config, evaluate fixture, and package files.
- Platform-specific setup patterns for evals-create-suite.
- Evidence-backed steps from upstream SKILL.md.
- When-to-use criteria for evals-create-suite versus alternatives.
Evals Create Suite by the numbers
- 2 all-time installs (skills.sh)
- Ranked #1,788 of 2,203 Security skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
evals-create-suite capabilities & compatibility
- Capabilities
- evals create suite quick start · evals create suite when to use guidance · evals create suite integration patterns
- Works with
- elasticsearch
- Use cases
- security audit
What evals-create-suite says it does
disable-model-invocation: true
**Suite name** (kebab-case, e.g. `my-feature`)
npx skills add https://github.com/elastic/kibana --skill evals-create-suiteAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 2 |
|---|---|
| repo stars | ★ 21.2k |
| Last updated | August 5, 2026 |
| Repository | elastic/kibana ↗ |
How do I use evals-create-suite correctly?
Scaffold a new LLM evaluation suite package with Playwright config, evaluate fixture, and package files. Use when creating a new eval suite, adding an evals package for a plugin, or setting up the boi
Who is it for?
Teams implementing evals-create-suite workflows from the catalog.
Skip if: Skip when requirements clearly match a different specialized stack.
When should I use this skill?
User asks about evals-create-suite, scaffold a new llm evaluation suite package with playwright config, evaluate fixture, and .
What you get
Working evals-create-suite setup with validated configuration and next steps.
Files
Create an Eval Suite
Overview
Eval suites live in dedicated kbn-evals-suite-<name> packages. Each suite is a self-contained Playwright project that uses the evaluate fixture from @kbn/evals to run LLM experiments with datasets, tasks, and evaluators.
Inputs to Collect
- Suite name (kebab-case, e.g.
my-feature) - Parent directory under
x-pack/(e.g.x-pack/platform/packages/shared/ai-infra/orx-pack/solutions/security/test/) - Owner GitHub team handle (e.g.
@elastic/appex-ai-infra) - Group (
platform,security,observability,search) - Visibility (
sharedorprivate) - Whether custom fixtures are needed (chat client, esArchiver, supertest, etc.)
Do NOT Use node scripts/scout.js generate
Eval suites are not standard Scout test configs. The Scout generator creates test/scout/ directories that are picked up by Scout's CI discovery glob -- this will break because evals configs use createPlaywrightEvalsConfig (not createPlaywrightConfig) and contain non-JS files (like .text prompt files) that Playwright cannot parse.
The Scout team has explicitly asked that eval configs live outside test/scout/ directories. All eval suites place their playwright.config.ts in the package root.
Directory Layout
kbn-evals-suite-<name>/
├── evals/
│ └── <name>.spec.ts # evaluation spec(s)
├── src/
│ └── evaluate.ts # re-export or extend the base evaluate fixture
├── playwright.config.ts # MUST be in package root, NOT under test/scout/
├── package.json
├── kibana.jsonc
└── tsconfig.jsonFile Templates
kibana.jsonc
{
"type": "functional-tests",
"id": "@kbn/evals-suite-<name>",
"owner": "@elastic/<team>",
"group": "<platform|security|observability|search>",
"visibility": "<shared|private>"
}type must be "functional-tests" -- not "shared-common" or "plugin".
package.json
{
"name": "@kbn/evals-suite-<name>",
"private": true,
"version": "1.0.0",
"license": "Elastic License 2.0"
}tsconfig.json
{
"extends": "@kbn/tsconfig-base/tsconfig.json",
"compilerOptions": {
"outDir": "target/types",
"types": ["jest", "node"]
},
"include": ["**/*.ts"],
"exclude": ["target/**/*"],
"kbn_references": [
"@kbn/evals",
"@kbn/scout"
]
}Add any additional package refs your suite imports to kbn_references (e.g. @kbn/inference-common, @kbn/es-archiver).
playwright.config.ts
import Path from 'path';
import { createPlaywrightEvalsConfig } from '@kbn/evals';
export default createPlaywrightEvalsConfig({
testDir: Path.resolve(__dirname, './evals'),
timeout: 30 * 60_000,
});Options:
testDir(required) -- directory containing.spec.tsfilestimeout(optional, default5 * 60_000) -- per-test timeout in msrepetitions(optional, default1) -- overridable viaEVALUATION_REPETITIONSenv var
src/evaluate.ts
Simple (no custom fixtures):
import { evaluate } from '@kbn/evals';
export { evaluate };Extended (with custom fixtures):
import { evaluate as base } from '@kbn/evals';
import { MyChatClient } from './chat_client';
export const evaluate = base.extend<
{},
{ chatClient: MyChatClient }
>({
chatClient: [
async ({ fetch, log, connector }, use) => {
await use(new MyChatClient(fetch, log, connector.id));
},
{ scope: 'worker' },
],
});When to Extend evaluate
Use the base evaluate directly when your task calls Kibana APIs through the built-in fetch, inferenceClient, or executorClient fixtures.
Extend when you need:
- A chat client that wraps a specific Kibana API endpoint (e.g.
/api/agent_builder/converse) - An `evaluateDataset` helper that encapsulates the
runExperiment+ evaluator wiring for a consistent pattern across specs - `esArchiver` for loading/unloading ES archives in setup/teardown
- `supertest` for direct HTTP assertions against Kibana
- Domain-specific API clients (e.g.
QuickstartClient)
Real examples
| Suite | Approach | Why |
|---|---|---|
llm-tasks | Base evaluate directly | Calls task functions in-process; custom CODE evaluators inline |
agent-builder | Extended with chatClient + Phoenix executor | Needs HTTP chat client and external Phoenix executor |
security-solution-evals | Extended with chatClient, esArchiver, supertest, quickApiClient | Domain-heavy setup: loads ES archives, uses generated API client |
Suite Registration
Add an entry to .buildkite/pipelines/evals/evals.suites.json:
{
"id": "<name>",
"name": "<Human Readable Name>",
"configPath": "<repo-relative path to playwright.config.ts>",
"tags": ["<group>", "<name>"],
"ciLabels": ["evals:<name>"]
}Registration is optional for local dev (suites are auto-discovered from createPlaywrightEvalsConfig imports), but required for CI labeling and node scripts/evals list.
Post-Scaffold Steps
1. Run yarn kbn bootstrap to register the new package. 2. Verify the suite appears: node scripts/evals list. 3. Create your first spec file under evals/ (see the evals-write-spec skill). 4. Run locally: node scripts/evals start --model <connector-id> --judge <connector-id>.
Common Mistakes
- Placing configs under `test/scout/` -- Scout's CI discovery will find them and crash. Keep
playwright.config.tsin the package root. - Using `node scripts/scout.js generate` -- this creates Scout test scaffolds, not eval suites. Scaffold manually using the templates above.
- Setting
typeto anything other than"functional-tests"inkibana.jsonc. - Forgetting
@kbn/evalsinkbn_references-- causes TS resolution failures. - Using
Path.joininstead ofPath.resolvefortestDir-- Playwright needs an absolute path. - Creating
evals/specs that import from@kbn/evalsbut the suite'ssrc/evaluate.tsre-exports a different fixture -- always importevaluatefrom the suite's ownsrc/evaluatewhen extending. - Forgetting to run
yarn kbn bootstrapafter creating the package.
Related skills
FAQ
What does evals-create-suite do?
evals-create-suite skill documents Scaffold a new LLM evaluation suite package with Playwright config, evaluate fixture, and package files.
When should I use evals-create-suite?
User asks about evals-create-suite, scaffold a new llm evaluation suite package with playwright config, evaluate fixture, and .
Is this skill safe to install?
Review the Security Audits panel on this page before installing in production.