
Test Smell Detection
- 9 installs
- 466 repo stars
- Updated July 25, 2026
- managedcode/dotnet-skills
Helps with testing & qa tasks.
About
test-smell-detection is a Claude Code skill for testing & qa. It helps solo builders move faster with AI-assisted coding.
- test-smell-detection
- Testing & QA
- AI-coding skill
Test Smell Detection by the numbers
- 9 all-time installs (skills.sh)
- +1 installs in the week ending Aug 2, 2026 (Skillselion tracking)
- Ranked #1,560 of 2,153 Testing & QA skills by installs in the Skillselion catalog
- Data as of Aug 3, 2026 (Skillselion catalog sync)
npx skills add https://github.com/managedcode/dotnet-skills --skill test-smell-detectionAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 9 |
|---|---|
| repo stars | ★ 466 |
| Last updated | July 25, 2026 |
| Repository | managedcode/dotnet-skills ↗ |
What it does
Helps with testing & qa tasks.
Files
Test Smell Detection
Deep formal audit of test code in any supported language using an academic test smell taxonomy. Detects symptoms of bad design or implementation decisions that make tests harder to understand, more fragile, less effective at catching bugs, or more expensive to maintain. Produces a severity-ranked report with specific locations and actionable fixes.
Language-specific guidance: Call the test-analysis-extensions skill to discover available extension files, then read the file matching the target codebase. The extension file documents test markers, sleep / time / random APIs, skip annotations, setup/teardown, mystery-guest indicators (file/database/network/env), integration markers, and language-specific calibration notes that drive the smell detectors below.Why Test Smells Matter
Test smells erode confidence in a test suite and inflate maintenance costs:
| Problem | Consequence |
|---|---|
| Tests with conditional logic | Some paths never execute — hidden testing gaps |
| Tests that depend on external resources | Flaky failures, slow execution, environment coupling |
| Tests that sleep to wait for results | Non-deterministic timing, slow suites, false failures |
| Tests without assertions | False confidence — coverage looks good but nothing is verified |
| Tests that call many production methods | Hard to diagnose failures, unclear what's being tested |
| Tests with magic numbers | Unreadable intent, unclear boundary conditions |
| Tests relying on ToString for comparison | Brittle to formatting changes, obscure failure messages |
| Tests with exception handling logic | Swallowed failures, tests that pass when they shouldn't |
When to Use
- User asks for a comprehensive or formal test smell audit
- User asks "are my tests well-written?" and wants a thorough analysis
- User wants a test quality health check with academic rigor
- User asks for a review of test design or structure using standard smell categories
- User suspects tests are fragile, flaky, or giving false confidence and wants a deep investigation
When Not to Use
- User wants a quick pragmatic test review (use
test-anti-patterns— faster, covers the most common issues) - User wants to evaluate assertion diversity specifically (use
assertion-quality) - User wants to find duplicated boilerplate across tests (use
exp-test-maintainability) - User wants to write new tests from scratch (help them directly)
- User wants to fix a specific failing test (diagnose and fix directly)
Inputs
| Input | Required | Description |
|---|---|---|
| Test code | Yes | One or more test files or a test project directory to analyze |
| Production code | No | The code under test, for context on whether patterns are justified |
Workflow
Step 1: Detect language and load extension
Identify the target codebase's language and test framework. Call the test-analysis-extensions skill and read the matching extension file (e.g., extensions/dotnet.md, extensions/python.md, extensions/typescript.md, extensions/go.md). The extension file lists the framework-specific test markers, sleep / wait APIs, skip / ignore attributes, mystery-guest indicators, and integration-test markers that the smell detectors below need.
Step 2: Gather the test code
Read all test files the user provides. If the user points to a directory or project, scan for all test files using the markers in the loaded language extension file.
For a thorough audit, also consult the extended smell catalog which covers 9 additional smell types beyond the core 10 below.
Step 3: Scan for test smells
For each test method and class, check for the following smell categories. Examples reference .NET attributes but the patterns apply across all supported languages — use the loaded language extension file to map each pattern to the framework you are auditing.
Smell 1: Conditional Test Logic
Test methods containing if, else, switch, ternary (? :), for, foreach, while, or pattern-match arms that change assertion behavior. Control flow in tests means some paths may never execute, hiding gaps.
Severity: High Detection: Any control-flow statement inside a test method body that affects which assertions run. Exceptions (per-language idioms, do NOT flag):
- Foreach-assert used solely to assert every item in a known collection (the assertion is the loop body).
- Go / Rust table-driven tests:
for _, tt := range tests { t.Run(tt.name, func(t *testing.T) { ... }) }(Go) or#[rstest]parametrized loops are idiomatic. - `it.each(...)` / `test.each(...)` / `@pytest.mark.parametrize` / `[Theory] + [InlineData]` / `@ParameterizedTest` parametrization driven by data tables.
- Pester `-ForEach` / `-TestCases` and RSpec `where` blocks.
- Catch2 `SECTION`s and `GENERATE(...)`, doctest `SUBCASE`, GoogleTest `INSTANTIATE_TEST_SUITE_P`.
Smell 2: Mystery Guest
Tests that depend on external resources — files on disk, databases, network endpoints, environment variables — without making the dependency explicit or using test doubles.
Severity: High Detection: Test methods that read files, open database connections, make HTTP requests (without a test handler), read environment variables, or use hard-coded file paths. Per language: File.ReadAllText / Directory.GetFiles / HttpClient / Environment.GetEnvironmentVariable (.NET); open() / pathlib.Path.read_text() / requests.get() / os.environ[...] (Python); fs.readFileSync / fetch(...) / process.env.X (JS/TS); Files.readAllBytes / Files.newInputStream / HttpClient.send / System.getenv (Java); os.ReadFile / http.Get / os.Getenv (Go); File.read / Net::HTTP.get / ENV[...] (Ruby); std::fs::read_to_string / reqwest::get / std::env::var (Rust); String(contentsOfFile:) / URLSession.shared.data / ProcessInfo.processInfo.environment (Swift); File(...).readText() / URL(...).openConnection() / System.getenv (Kotlin); Get-Content / Invoke-WebRequest / $env:X (Pester); std::ifstream / curl_easy_perform / std::getenv (C++). Exception: In-memory fakes, test-specific handlers, or hermetic test data factories are fine.
Smell 3: Sleepy Test
Tests that call sleep or delay functions to wait for a condition. These introduce non-deterministic timing and slow down the suite.
Severity: High Detection: Calls to sleep/delay functions inside test methods: Thread.Sleep / Task.Delay (.NET); time.sleep / asyncio.sleep (Python); setTimeout / await new Promise(r => setTimeout(...)) / jest.advanceTimersByTime not paired with a wait (JS/TS); Thread.sleep / TimeUnit.SECONDS.sleep (Java); time.Sleep (Go); sleep / Kernel#sleep (Ruby); std::thread::sleep / tokio::time::sleep (Rust); Thread.sleep / delay (Kotlin coroutines); sleep(_:) / Task.sleep (Swift); Start-Sleep (Pester); std::this_thread::sleep_for (C++). See the matching language extension file for the full list.
Smell 4: Assertion-Free Test (Unknown Test)
Tests that execute code but never assert anything. Test frameworks report these as passing even if the code is completely broken, as long as no exception is thrown.
Severity: High Detection: A test method with no assertion calls and no expected-exception annotation. Framework-specific: missing Assert.* (.NET); no assert / pytest.raises (Python); no expect(...) or assert.* (JS/TS); no assert* / assertThat (Java); no t.Error* / t.Fatal* / assert.* testify (Go); no expect/.to/.eq (RSpec) or assert*/refute* (Minitest); no assert*! / assert_eq! / panic! (Rust); no XCTAssert* / #expect (Swift); no assert* / should* / Kotest matchers (Kotlin); no Should -* (Pester); no EXPECT_* / ASSERT_* / REQUIRE / CHECK (C++). Calibration:
- A method named
*_DoesNotThrow/*_no_exception/should not throwis implicitly asserting no exception — still flag it but note it may be intentional. - Mock-call verifications count as assertions:
mock.Verify(...)(Moq),Mock.AssertWasCalled(NSubstitute),mock.assert_called_with(...)(Python),expect(mock).toHaveBeenCalledWith(...)(Jest),verify(mock).method(...)(Mockito),Should -Invoke(Pester) — do NOT flag tests using these as assertion-free. - Bare assertion forms count:
assert x == y(pytest),if got != want { t.Errorf(...) }(Go),assert!(cond)(Rust) are canonical. - Snapshot assertions count:
.toMatchSnapshot()(Jest),syrupy(pytest),SnapshotTesting(Swift),approval-testsare real assertions. - Missing await on async assertions is its own critical smell:
expect(promise).resolves.toBe(x)withoutawait/return(Jest), un-awaitedAssert.ThrowsAsync(xUnit), un-awaited coroutines inpytest-asyncio, Kotest tests withoutrunTest, Swift Testing async cases withoutawait. These tests have assertion calls but silently pass — flag with a dedicated note.
Smell 5: Eager Test
A test method that calls many different production methods, making it unclear what behavior is being tested. When it fails, diagnosis is difficult because the failure could stem from any of the calls.
Severity: Medium Detection: A test method that calls 4+ distinct methods on the production object (excluding setup/construction). Count unique method names, not call count. Calibration: Integration / end-to-end / workflow tests may legitimately call multiple methods. Check for integration markers in the loaded language extension file (e.g., [Trait("Category", "Integration")], @Tag("integration"), pytest.mark.integration, *_integration_test.go, Describe ... -Tag 'Integration') and downgrade.
Smell 6: Magic Number Test
Assertions that contain unexplained numeric literals. The intent of Assert.AreEqual(42, result) / assert result == 42 / expect(result).toBe(42) is unclear without context — what does 42 represent?
Severity: Medium Detection: Numeric literals (other than 0, 1, -1, and the literal used in the test name) appearing as expected parameters in assertion methods or comparison operands. Calibration: Small integers in context (like count checks Assert.AreEqual(3, list.Count) / assert len(items) == 3 / expect(arr.length).toBe(3) where 3 items were just added) are acceptable — only flag when the number's meaning is genuinely unclear.
Smell 7: Sensitive Equality
Tests that use string conversion for comparison or assertion. If the underlying string representation changes, the test breaks even though the actual behavior is correct.
Severity: Medium Detection: Assert.AreEqual(expected, obj.ToString()) (.NET); assert str(obj) == "..." or assert repr(obj) == "..." (Python); expect(obj.toString()).toBe("...") or expect(${obj}).toBe(...) (JS/TS); assertEquals(expected, obj.toString()) (Java); assert.Equal(t, "...", fmt.Sprint(obj)) or obj.String() chains (Go); expect(obj.to_s).to eq("...") (RSpec); assert_eq!(format!("{}", obj), "...") or assert_eq!(format!("{:?}", obj), "...") (Rust); XCTAssertEqual(obj.description, "...") or string-interpolation assertion (Swift); assertEquals("...", obj.toString()) (Kotlin); Should -Be "..." against a [string]$obj (Pester); EXPECT_EQ("...", std::to_string(obj)) (C++).
Smell 8: Exception Handling in Tests
Tests that contain try/catch/except/rescue blocks or throw/raise/panic/return err statements used to manage exception flow instead of asserting on it. This typically means the test is manually managing errors rather than using the framework's built-in exception assertion facilities.
Severity: Medium Detection: try/catch (.NET, Java, JS/TS, Kotlin, Swift, C++); try/except (Python); begin/rescue (Ruby); defer recover() (Go); manual if err != nil { t.Fatal(err) } in Go is canonical and NOT a smell. Exception: catch/except/rescue blocks that capture an exception for further assertion on its properties are a lesser concern — note but don't flag as high severity.
Smell 9: General Fixture (Over-broad Setup)
The test setup method, constructor, or fixture initializes fields that are not used by every test method. This means each test pays the cost of setting up objects it doesn't need.
Severity: Low Detection: Fields/properties initialized in [TestInitialize] / setUp / @BeforeEach / beforeEach / before(:each) / BeforeEach (Pester) / setUpWithError (XCTest) / pytest fixture(autouse=True) / xUnit constructor / Kotest beforeTest that are referenced by fewer than half the test methods in the class/module/file.
Smell 10: Ignored / Disabled / Skipped Test
Tests marked as skipped or disabled. These add overhead and clutter, and the underlying issue they were disabled for may never be addressed.
Severity: Low Detection: Skip / ignore / disable annotations or conditional compilation that disables a test. See the loaded language extension file for framework-specific skip attributes — e.g., [Ignore] (MSTest/NUnit), Skip = "..." (xUnit Fact), @Ignore (TUnit/JUnit 4), @Disabled (JUnit 5), @pytest.mark.skip / pytest.skip(...) / pytestmark, it.skip / xit / describe.skip / test.skip (Jest/Vitest/Mocha), t.Skip(...) (Go), pending / skip / xit (RSpec), #[ignore] (Rust), XCTSkip / @Test(.disabled) (Swift), @Ignored (Kotest), -Skip (Pester), GTEST_SKIP() / DISABLED_TestName (GoogleTest), [.] tag (Catch2), TEST_CASE("...", "[.]") skip.
Step 4: Apply calibration rules
Before reporting, calibrate findings to avoid false positives:
- Integration tests have different norms. A test class clearly marked as integration (by name, annotation, category, or convention — see the loaded language extension file for markers) legitimately uses external resources, calls multiple methods, and may use delays for async coordination. Downgrade Mystery Guest, Eager Test, and Sleepy Test severity for integration tests — note them but don't flag as problems.
- Simple loop-assert patterns are fine. Iterating a collection to assert on every item is readable and correct. Only flag loops with complex branching logic.
- Idiomatic table-driven and parametrized patterns are NOT Conditional Test Logic. Go's
for _, tt := range tests { t.Run(...) }, Rust's#[rstest], pytest's@parametrize, Jest/Vitest.each, JUnit@ParameterizedTest, RSpecwhere, Pester-ForEach, Catch2SECTION/GENERATE, GoogleTestINSTANTIATE_TEST_SUITE_Pare canonical and must NOT be flagged. - Context matters for magic numbers. A count assertion right after adding a known number of items is self-documenting. Only flag numbers whose meaning requires looking at production code to understand.
- Bare `assert` (pytest) is canonical, not assertion-free framework use. Don't flag.
- Go's `if err != nil { t.Fatal(err) }` is canonical, not Exception Handling in Tests. Don't flag.
- Mock-call verifications and snapshot assertions are real assertions — do not flag tests using them as Assertion-Free.
- Missing-await on async assertions is its own critical sub-smell of Assertion-Free — these tests silently pass even when the underlying assertion fails. Always flag when detected.
- Inconclusive/pending markers are not assertion-free. Tests explicitly marked as incomplete should be flagged as Ignored Test, not Assertion-Free.
- Capture-and-assert exception patterns are borderline.
try { ... } catch (X x) { Assert.Equal(...) }style patterns are ugly but functional. Note as a smell and suggest the framework's built-in exception assertion (Assert.Throws<T>,pytest.raises,expect(fn).toThrow,assertThrows,assert.PanicsWithError, etc.) instead of calling it broken. - If the test suite is clean, say so. A report finding few or no smells is perfectly valid.
Step 5: Report findings
Present the analysis in this structure:
1. Summary Dashboard — Quick overview:
| Severity | Smell Count | Affected Tests |
|----------|-------------|----------------|
| High | 3 | 7 |
| Medium | 2 | 4 |
| Low | 1 | 2 |
| Total | 6 | 13 |2. Findings by Severity — For each smell found:
- Smell name and category
- Severity level with rationale
- Affected test methods (file and method name)
- Code snippet showing the smell
- Concrete fix: show what the code should look like after remediation
- Risk if left unfixed
3. Smell-Free Patterns — If any test methods are well-written, briefly acknowledge this. Highlighting what's good helps the user understand the contrast.
4. Prioritized Remediation Plan — Rank fixes by:
- Impact (high-severity smells affecting many tests first)
- Effort (quick fixes before refactoring)
- Risk (fixes that prevent false-passes before cosmetic improvements)
Validation
- [ ] Every finding includes the specific test method name and file location
- [ ] Every finding includes a code snippet showing the smell in context
- [ ] Every finding includes a concrete fix example (not just "fix this")
- [ ] Integration tests are not penalized for patterns that are appropriate for their scope
- [ ] Simple foreach-assert loops are not flagged as conditional test logic
- [ ] Contextually obvious numbers are not flagged as magic numbers
- [ ] If the test suite is clean, the report says so upfront
- [ ] Severity levels are justified, not arbitrary
Common Pitfalls
| Pitfall | Solution |
|---|---|
| Flagging integration tests for using real resources | Check for integration test markers (per the loaded language extension) and adjust severity accordingly |
| Flagging loop-over-collection-assert as conditional logic | Only flag loops with branching or complex logic, not assertion iterations |
| Flagging Go/Rust table-driven loops as Conditional Test Logic | for _, tt := range tests { t.Run(...) } (Go) and #[rstest] loops (Rust) are canonical and must NOT be flagged |
| Flagging parametrized tests as Duplicate Assert | @pytest.mark.parametrize, it.each, [Theory]+[InlineData], @ParameterizedTest, RSpec where, Pester -ForEach, Catch2 SECTION/GENERATE are correct deduplication, not smells |
Flagging pytest bare assert as missing framework | Bare assert is canonical pytest assertion — count it |
Flagging Go's if err != nil { t.Fatal(err) } as Exception Handling in Tests | This is canonical Go error checking — do NOT flag |
| Flagging obvious count assertions after adding N items | Consider the immediate context — self-documenting numbers are fine |
| Missing framework-specific assertion syntax | Always read the matching language extension file first; each framework has distinct assertion APIs (xUnit Assert.Equal, MSTest Assert.AreEqual, NUnit Is.EqualTo, pytest bare assert, Jest expect().toBe(), etc.) |
| Treating mock-call verifications as assertion-free | mock.Verify(...), expect(mock).toHaveBeenCalledWith(...), Should -Invoke, verify(mock).method(...), mock.assert_called_with(...) are real assertions |
| Missing the async-test silent-pass trap | Always flag expect(promise).resolves.toBe(x) without await/return, un-awaited Assert.ThrowsAsync (xUnit), un-awaited coroutines in pytest-asyncio, missing runTest in Kotest, un-awaited Swift Testing async assertions |
| Over-flagging try/catch that captures for assertion | Distinguish swallowed exceptions from capture-and-assert patterns |
| Treating skip annotations with reasons same as bare skips | Note that reasoned skips (Skip = "Tracked by #123", @pytest.mark.skip(reason="..."), t.Skip("not yet implemented")) are less concerning than unexplained ones |
Flagging DoesNotThrow-style tests as assertion-free | These implicitly assert no exception — note but acknowledge the intent |
{
"version": "0.1.0",
"category": "Testing",
"compatibility": "Requires a .NET test project or solution."
}
Test Smell Catalog
Extended catalog of test smells based on academic research. This reference provides deeper background on smells beyond those covered in the core skill, including research origins, prevalence data, and real-world examples from open-source projects.
Source: testsmells.org — a research project from the Rochester Institute of Technology.
Full Smell Taxonomy
The academic literature identifies 19 distinct test smell types. The core skill covers the 10 most impactful ones. This catalog documents all 19 for deeper analysis when requested.
Smells Covered by the Core Skill
| Smell | Core Skill | Academic Name |
|---|---|---|
| Conditional Test Logic | Smell 1 | Conditional Test Logic |
| Mystery Guest | Smell 2 | Mystery Guest |
| Sleepy Test | Smell 3 | Sleepy Test |
| Assertion-Free Test | Smell 4 | Unknown Test / Empty Test |
| Eager Test | Smell 5 | Eager Test |
| Magic Number Test | Smell 6 | Magic Number Test |
| Sensitive Equality | Smell 7 | Sensitive Equality |
| Exception Handling | Smell 8 | Exception Handling |
| General Fixture | Smell 9 | General Fixture |
| Ignored/Disabled Test | Smell 10 | Ignored Test |
Additional Smells (Extended Analysis)
These smells are not in the core skill but can be reported when the user requests a thorough audit or when they are particularly prevalent.
Assertion Roulette
A test method has multiple assertions without descriptive messages. When one fails, it's unclear which assertion caused the failure and why.
Detection: A test method containing 3+ assertion statements where none provide an explanation message parameter.
Example:
assertThat(repo, hasGitObject("ba1f63e4430bff267d112b1e8afc1d6294db0ccc"));
File readmeFile = new File(repo.getWorkTree(), "README");
assertThat(readmeFile, exists());
assertThat(readmeFile, ofLength(12));Three assertions, no messages — if one fails, you must read the code to determine which property was wrong.
Duplicate Assert
The same assertion (same parameters) appears multiple times in a single test method.
Detection: Two or more assertion statements within the same test method with identical parameters.
Example:
valid = XmlSanitizer.isValid("Fritz-box");
assertEquals("Minus is valid", true, valid);
// ... later in the same method:
valid = XmlSanitizer.isValid("Fritz-box");
assertEquals("Minus is valid", true, valid);Lazy Test
Multiple test methods test the same production method. While not always a problem, it may indicate tests that should be parameterized or that explore the same behavior redundantly.
Detection: Multiple test methods in the same class calling the same production method as their primary action.
Constructor Initialization
Test class uses a constructor instead of the framework's setup method to initialize fields. This bypasses framework lifecycle hooks and can cause issues with test isolation.
Detection: Test class has a constructor that initializes fields rather than using the designated setup method.
Default Test
The test class retains its auto-generated template name (e.g., ExampleUnitTest, UnitTest1). This indicates the test file was scaffolded but never properly organized.
Detection: Test class named ExampleUnitTest, ExampleInstrumentedTest, UnitTest1, TestClass1, or similar template names.
Redundant Print
Test methods contain Console.WriteLine, System.out.println, print(), or similar output statements. These are debugging artifacts that add noise and slow execution.
Detection: Print/log statements inside test methods that are not part of a logging-focused test.
Redundant Assertion
Assertions that are always true or always false regardless of the code under test.
Detection: Assertions comparing a value to itself, or asserting literal true/false constants.
Example:
assertEquals(true, true);Resource Optimism
Tests that assume external resources (files, services) exist without checking. The test may pass locally but fail in CI or on another developer's machine.
Detection: File or resource references used without existence checks or guard assertions.
Empty Test
A test method that contains no executable statements — only comments or whitespace. Similar to Assertion-Free Test but even more extreme: no code runs at all.
Detection: Test method body contains only comments, whitespace, or commented-out code.
Research Background
Prevalence
Research on Android open-source projects (Peruma et al., CASCON 2019) found:
- Assertion Roulette and Eager Test are the most common smells
- Over 50% of test files contain at least one smell
- Smells tend to accumulate over time — they are rarely refactored away
Impact on Flakiness
Camara et al. (SAST 2021) found a correlation between test smells and flaky tests — smelly tests are more likely to exhibit non-deterministic failures.
Severity Thresholds
Spadini et al. (MSR 2020) investigated severity thresholds for test smells, finding that developer perception of smell severity varies significantly. Some smells (Conditional Test Logic, Sleepy Test) are consistently rated as serious, while others (Magic Number, Assertion Roulette) are considered minor annoyances.
Key Publications
- Peruma et al. (2020). "tsDetect: An Open Source Test Smells Detection Tool." ESEC/FSE 2020.
- Peruma et al. (2019). "On the Distribution of Test Smells in Open Source Android Applications." CASCON 2019.
- Spadini et al. (2020). "Investigating Severity Thresholds for Test Smells." MSR 2020.
- Camara et al. (2021). "On the use of test smells for prediction of flaky tests." SAST 2021.
- Kim et al. (2021). "The secret life of test smells — an empirical study on test smell evolution and maintenance." Empirical Software Engineering, 26(100).