Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
mastra-ai avatar

E2e Tests Studio

  • 1.6k installs
  • 26.9k repo stars
  • Updated August 5, 2026
  • mastra-ai/mastra

e2e-tests-studio is an agent skill for writing Playwright behavior E2E tests for Mastra playground UI with required BDD structure.

About

The e2e-tests-studio skill generates Playwright E2E tests for packages/playground-ui and packages/playground that validate product behavior not UI states. Required BDD structure uses one outer test.describe for the unit, inner test.describe when preconditions, and each test asserting one observable outcome enforced by e2e-bdd/test-needs-when-describe. Behavior patterns cover configuration affecting agent responses, data persistence after reload, tool execution output content, workflow step chaining, streaming chat context, and error recovery retries. Kitchen-sink fixtures represent realistic scenarios with expectedBehavior metadata. Use when modifying React components, playground features, or studio UI requiring behavior-focused Playwright specs. Agents should follow the SKILL.md workflow end to end, grounding classification in documented commands, file paths, prerequisites, and troubleshooting notes rather than improvising steps. Write Playwright E2E behavior tests for Mastra playground UI changes with required BDD describe structure. Invoke when User modifies playground UI, needs E2E tests, or mentions Playwright behavior validation. Best for Developers modifying Mastra playgrou.

  • Test product behavior not UI states like dropdown opens.
  • Required BDD: outer unit describe, when precondition, one outcome per test.
  • Behavior patterns: config affects API, persistence, tool output, workflows.
  • Kitchen-sink fixtures with scenario description and expectedBehavior.
  • Run via pnpm build:cli then kitchen-sink dev on localhost:4111.

E2e Tests Studio by the numbers

  • 1,602 all-time installs (skills.sh)
  • +77 installs in the week ending Aug 4, 2026 (Skillselion tracking)
  • Ranked #467 of 2,153 Testing & QA skills by installs in the Skillselion catalog
  • Security screen: LOW risk (skills.sh audit)
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
At a glance

e2e-tests-studio capabilities & compatibility

Capabilities
bdd test structure enforcement · behavior focused test patterns · kitchen sink fixture design · playwright persistence and api interception test
Use cases
testing · frontend
From the docs

What e2e-tests-studio says it does

Tests must verify that product features WORK correctly, not just that UI elements render.
SKILL.md
Every E2E spec MUST follow the same BDD shape as the MSW tests.
SKILL.md
npx skills add https://github.com/mastra-ai/mastra --skill e2e-tests-studio

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs1.6k
repo stars26.9k
Security audit3 / 3 scanners passed
Last updatedAugust 5, 2026
Repositorymastra-ai/mastra

How do I write Playwright E2E tests that verify Mastra playground product behavior not just UI states?

Write Playwright E2E behavior tests for Mastra playground UI changes with required BDD describe structure.

Who is it for?

Developers modifying Mastra playground or playground-ui needing behavior validation tests.

Skip if: Skip for unit tests, backend-only changes, or UI snapshot tests without behavior assertions.

When should I use this skill?

User modifies playground UI, needs E2E tests, or mentions Playwright behavior validation.

What you get

BDD-structured Playwright specs asserting observable outcomes for playground feature changes.

  • Playwright E2E spec files
  • Behavior-focused test scenarios

By the numbers

  • Targets two Mastra packages: packages/playground-ui and packages/playground
  • Specifies model claude-opus-4-5 in skill metadata

Files

SKILL.mdMarkdownGitHub ↗

E2E Behavior Validation for Frontend Modifications

Core Principle: Test Product Behavior, Not UI States

CRITICAL: Tests must verify that product features WORK correctly, not just that UI elements render.

What NOT to test (UI States):

  • ❌ "Dropdown opens when clicked"
  • ❌ "Modal appears after button click"
  • ❌ "Loading spinner shows during request"
  • ❌ "Form fields are visible"
  • ❌ "Sidebar collapses"

What TO test (Product Behavior):

  • ✅ "Selecting an LLM provider configures the agent to use that provider"
  • ✅ "Creating a new agent persists it and shows in the agents list"
  • ✅ "Running a tool with parameters returns the expected output"
  • ✅ "Chat messages stream correctly and maintain conversation context"
  • ✅ "Workflow execution triggers tools in the correct order"

BDD Structure (REQUIRED)

Every E2E spec MUST follow the same BDD shape as the MSW tests. In packages/playground, e2e-bdd/test-needs-when-describe enforces this shape.

The structure has exactly three levels:

1. Outer `test.describe` = the unit under test (one page or feature per file). 2. Inner `test.describe('when …')` = exactly ONE precondition. The title MUST start with when. 3. Each `test` = exactly ONE observable outcome.

import { test, expect } from '@playwright/test';
import { resetStorage } from '../__utils__/reset-storage';

test.describe('Tools list page', () => {
  // the unit
  test.afterEach(async () => {
    await resetStorage();
  });

  test.describe('when a registered tool is clicked', () => {
    // ONE precondition (starts with "when")
    test('navigates to that tool detail page', async ({ page }) => {
      // ONE outcome
      await page.goto('/tools');
      await page.locator('text=Get current weather for a location').click();
      await expect(page).toHaveURL(/\/tools\/weatherInfo$/);
    });

    test('shows the tool name as the page heading', async ({ page }) => {
      // ONE outcome
      await page.goto('/tools');
      await page.locator('text=Get current weather for a location').click();
      await expect(page.locator('h2')).toHaveText('weatherInfo');
    });
  });
});

Rules:

  • One outer test.describe per file naming the unit.
  • Every leaf test lives inside a test.describe('when …') precondition group. No top-level flat `test()`.
  • Split a multi-assertion test() only where assertions represent distinct outcomes; keep tightly-coupled assertions that prove a single outcome together. Never drop an assertion.
  • Place beforeEach/afterEach in the narrowest describe scope that needs them.

Prerequisites

Requires Playwright MCP server. If the browser_navigate tool is unavailable, instruct the user to add it:

claude mcp add playwright -- npx @playwright/mcp@latest

Step 1: Understand the Feature Intent

Before writing ANY test, answer these questions:

1. What user problem does this feature solve? 2. What is the expected outcome when the feature works correctly? 3. What data flows through the system? (user input → API → state → UI) 4. What should persist after page reload? 5. What downstream effects should this action have?

Document these answers as comments in your test file.

Step 2: Build and Start

pnpm build:cli
cd packages/playground/e2e/kitchen-sink && pnpm dev

Verify server at http://localhost:4111

Step 3: Map Feature to Behavior Tests

Feature-to-Test Mapping Guide

Feature CategoryWhat to TestExample Assertion
Agent ConfigurationConfig changes affect agent behaviorSend message → verify response uses selected model
LLM Provider SelectionSelected provider is used in requestsIntercept API call → verify provider in request payload
Tool ExecutionTool runs with correct params & returns resultExecute tool → verify output matches expected transformation
Workflow ExecutionSteps execute in order, data flows between stepsRun workflow → verify each step's output feeds next step
Chat/StreamingMessages persist, context maintained across turnsMulti-turn conversation → verify context awareness
MCP Server ToolsServer tools are callable and return dataCall MCP tool → verify response structure and content
Memory/PersistenceData survives page reloadCreate item → reload → verify item exists
Error HandlingErrors surface correctly to userTrigger error condition → verify error message + recovery

Step 4: Write Behavior-Focused Tests

Test Structure Template

import { test, expect, Page } from '@playwright/test';
import { resetStorage } from '../__utils__/reset-storage';
import { selectFixture } from '../__utils__/select-fixture';
import { nanoid } from 'nanoid';

/**
 * FEATURE: [Name of feature]
 * USER STORY: As a user, I want to [action] so that [outcome]
 * BEHAVIOR UNDER TEST: [Specific behavior being validated]
 */

test.describe('[Feature Name] - Behavior Tests', () => {
  let page: Page;

  test.beforeEach(async ({ browser }) => {
    const context = await browser.newContext();
    page = await context.newPage();
  });

  test.afterEach(async () => {
    await resetStorage(page);
  });

  test.describe('when [the single precondition for these outcomes]', () => {
    test('[verb describing the single observable outcome]', async () => {
      // ARRANGE: Set up preconditions
      // - Navigate to the feature
      // - Configure any required state
      // ACT: Perform the user action that triggers the behavior
      // ASSERT: Verify the OUTCOME, not the UI state
      // - Check data persistence
      // - Verify downstream effects
      // - Confirm API calls made correctly
    });
  });
});

Behavior Test Patterns

Pattern 1: Configuration Affects Behavior
test.describe('when a different LLM provider is selected', () => {
  test('uses that provider for agent responses', async () => {
    // ARRANGE
    await page.goto('/agents/my-agent/chat');

    // Intercept API to verify provider
    let capturedProvider: string | null = null;
    await page.route('**/api/chat', route => {
      const body = JSON.parse(route.request().postData() || '{}');
      capturedProvider = body.provider;
      route.continue();
    });

    // ACT: Select a different provider
    await page.getByTestId('provider-selector').click();
    await page.getByRole('option', { name: 'OpenAI' }).click();

    // Send a message to trigger the agent
    await page.getByTestId('chat-input').fill('Hello');
    await page.getByTestId('send-button').click();

    // ASSERT: Verify the selected provider was used
    await expect.poll(() => capturedProvider).toBe('openai');
  });
});
Pattern 2: Data Persistence
test.describe('when a new agent is created', () => {
  test('persists after page reload', async () => {
    // ARRANGE
    await page.goto('/agents');
    const agentName = `Test Agent ${nanoid()}`;

    // ACT: Create new agent
    await page.getByTestId('create-agent-button').click();
    await page.getByTestId('agent-name-input').fill(agentName);
    await page.getByTestId('save-agent-button').click();

    // Wait for creation to complete
    await expect(page.getByText(agentName)).toBeVisible();

    // ASSERT: Verify persistence
    await page.reload();
    await expect(page.getByText(agentName)).toBeVisible({ timeout: 10000 });
  });
});
Pattern 3: Tool Execution Produces Correct Output
test.describe('when the weather tool is executed with a city', () => {
  test('returns formatted weather data for that city', async () => {
    // ARRANGE
    await selectFixture(page, 'weather-success');
    await page.goto('/tools/weather-tool');

    // ACT: Execute tool with parameters
    await page.getByTestId('param-city').fill('San Francisco');
    await page.getByTestId('execute-tool-button').click();

    // ASSERT: Verify OUTPUT content, not just that output appears
    const output = page.getByTestId('tool-output');
    await expect(output).toContainText('temperature');
    await expect(output).toContainText('San Francisco');

    // Verify structured data if applicable
    const outputText = await output.textContent();
    const outputData = JSON.parse(outputText || '{}');
    expect(outputData).toHaveProperty('temperature');
    expect(outputData).toHaveProperty('conditions');
  });
});
Pattern 4: Workflow Step Chaining
test.describe('when a multi-step workflow is run', () => {
  test('passes data between steps correctly', async () => {
    // ARRANGE
    await selectFixture(page, 'workflow-multi-step');
    const sessionId = nanoid();
    await page.goto(`/workflows/data-pipeline?session=${sessionId}`);

    // ACT: Trigger workflow execution
    await page.getByTestId('workflow-input').fill('test input data');
    await page.getByTestId('run-workflow-button').click();

    // ASSERT: Verify each step received correct input from previous step
    // Wait for completion
    await expect(page.getByTestId('workflow-status')).toHaveText('completed', { timeout: 30000 });

    // Check step outputs show data transformation chain
    const step1Output = await page.getByTestId('step-1-output').textContent();
    const step2Output = await page.getByTestId('step-2-output').textContent();

    // Verify step 2 received step 1's output as input
    expect(step2Output).toContain(step1Output);
  });
});
Pattern 5: Streaming Chat with Context
test.describe('when a multi-turn conversation is held', () => {
  test('maintains conversation context across messages', async () => {
    // ARRANGE
    await selectFixture(page, 'contextual-chat');
    const chatId = nanoid();
    await page.goto(`/agents/assistant/chat/${chatId}`);

    // ACT: Multi-turn conversation
    await page.getByTestId('chat-input').fill('My name is Alice');
    await page.getByTestId('send-button').click();
    await expect(page.getByTestId('assistant-message').last()).toBeVisible({ timeout: 20000 });

    await page.getByTestId('chat-input').fill('What is my name?');
    await page.getByTestId('send-button').click();

    // ASSERT: Verify context was maintained
    const response = page.getByTestId('assistant-message').last();
    await expect(response).toContainText('Alice', { timeout: 20000 });
  });
});
Pattern 6: Error Recovery
test.describe('when the API fails during tool execution', () => {
  test('shows an actionable error and allows a successful retry', async () => {
    // ARRANGE: Set up failure fixture
    await selectFixture(page, 'api-failure');
    await page.goto('/tools/flaky-tool');

    // ACT: Trigger the error
    await page.getByTestId('execute-tool-button').click();

    // ASSERT: Error is shown with recovery option
    await expect(page.getByTestId('error-message')).toContainText('failed');
    await expect(page.getByTestId('retry-button')).toBeVisible();

    // Switch to success fixture and retry
    await selectFixture(page, 'api-success');
    await page.getByTestId('retry-button').click();

    // Verify recovery worked
    await expect(page.getByTestId('tool-output')).toBeVisible({ timeout: 10000 });
    await expect(page.getByTestId('error-message')).not.toBeVisible();
  });
});

Step 5: Update Existing Tests

When a test file already exists:

1. Read the existing tests to understand current coverage 2. Identify if tests are UI-focused or behavior-focused 3. Refactor UI-focused tests to verify behavior instead:

Refactoring Example

BEFORE (UI-focused):

test('dropdown opens when clicked', async () => {
  await page.getByTestId('model-dropdown').click();
  await expect(page.getByRole('listbox')).toBeVisible();
});

AFTER (Behavior-focused + BDD nesting):

test.describe('when a model is selected from the dropdown', () => {
  test('updates and persists the agent configuration', async () => {
    // Open dropdown and select model
    await page.getByTestId('model-dropdown').click();
    await page.getByRole('option', { name: 'GPT-4' }).click();

    // Verify the selection persists and affects behavior
    await page.reload();
    await expect(page.getByTestId('model-dropdown')).toHaveText('GPT-4');

    // Optionally: verify the model is used in actual requests
    // (via request interception or checking response metadata)
  });
});

Step 6: Kitchen-Sink Fixtures for Behavior Testing

Fixtures should represent realistic scenarios, not just mock data:

Fixture Naming Convention

<feature>-<scenario>.fixture.ts

Examples:
- agent-with-tools.fixture.ts
- chat-multi-turn-context.fixture.ts
- workflow-parallel-execution.fixture.ts
- tool-validation-error.fixture.ts
- mcp-server-timeout.fixture.ts

Fixture Content Requirements

Each fixture must define:

1. Scenario description (what behavior it enables testing) 2. Expected outcomes (what assertions should pass) 3. Edge cases covered (error states, empty states, etc.)

// fixtures/agent-provider-switch.fixture.ts
export const agentProviderSwitch = {
  name: 'agent-provider-switch',
  description: 'Tests that switching LLM providers changes agent behavior',

  // Mock responses for different providers
  responses: {
    openai: { content: 'Response from OpenAI', model: 'gpt-4' },
    anthropic: { content: 'Response from Anthropic', model: 'claude-3' },
  },

  expectedBehavior: {
    // When provider is switched, subsequent messages use new provider
    providerSwitchAffectsNextMessage: true,
    // Provider selection persists across page reload
    providerPersistsOnReload: true,
  },
};

Step 7: Run and Validate

cd packages/playground && pnpm test:e2e

Test Quality Checklist

Before considering tests complete, verify:

  • [ ] Each test has a clear user story comment
  • [ ] One outer test.describe names the unit under test
  • [ ] Every test is nested in a test.describe('when …') precondition block (no flat top-level test())
  • [ ] Each test asserts exactly ONE observable outcome
  • [ ] Tests verify OUTCOMES, not intermediate UI states
  • [ ] Tests would FAIL if the feature broke (not just if UI changed)
  • [ ] Persistence is verified via page.reload() where applicable
  • [ ] Error scenarios are covered
  • [ ] Tests use appropriate timeouts for async operations
  • [ ] Fixtures represent realistic usage scenarios

Quick Reference

StepCommand/Action
Buildpnpm build:cli
Startcd packages/playground/e2e/kitchen-sink && pnpm dev
App URLhttp://localhost:4111
Routes@packages/playground/src/App.tsx
Run testscd packages/playground && pnpm test:e2e
Test dirpackages/playground/e2e/tests/
Fixturespackages/playground/e2e/kitchen-sink/fixtures/

Anti-Patterns to Avoid

❌ Don't✅ Do Instead
Test that modal opensTest that modal action completes and persists
Test that button is clickableTest that clicking button produces expected result
Test loading spinner appearsTest that loaded data is correct
Test form validation message showsTest that invalid form cannot submit AND valid form succeeds
Test dropdown has optionsTest that selecting option changes system behavior
Test sidebar navigation worksTest that navigated page has correct data/functionality
Assert element is visibleAssert element contains expected data/state
Top-level flat test() with no precondition describeNest every test in a test.describe('when …') block
One test() asserting several unrelated outcomesOne test() per observable outcome

Related skills

How it compares

Use this instead of generic React testing skills when Mastra playground changes need behavior-focused Playwright coverage with documented UI-state anti-patterns.

FAQ

What must every E2E spec follow?

One outer describe for the unit, inner describe('when …') preconditions, and each test asserting one observable outcome.

UI states or behavior?

Behavior — verify configuration changes agent responses, data persists, tools return correct output, not that modals open.

Is e2e-tests-studio safe to install?

Review the Security Audits panel on this page before installing in production.

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.