
Symbolic Equation
- 1.2k installs
- 255 repo stars
- Updated February 27, 2026
- lingzhi227/agent-research-skills
symbolic-equation is a Claude Code skill that discovers symbolic mathematical equations for scientific systems using LLM-driven multi-island evolutionary search for developers who need interpretable formulas from data.
About
symbolic-equation is a lingzhi227/agent-research-skills agent skill extracted from the LLM-SR codebase spanning pipeline.py, sampler.py, buffer.py, evaluator.py, and config.py. It runs a multi-island evolutionary pipeline where parallel LLM samplers propose equation function structures, an ExperienceBuffer clusters programs across islands, and evaluators score fitness against scientific inputs. Prompts ask the model to complete equation functions respecting physical meaning and variable relationships. Developers reach for symbolic-equation when fitting compact symbolic models to experimental datasets instead of opaque black-box neural regressors.
- Multi-island evolutionary algorithm with parallel LLM samplers and evaluators
- ExperienceBuffer maintains versioned improving program clusters across islands
- Automated prompt construction from physical system relationships
- Continuous sampling loop with early stopping based on global sample count
- Parallel evaluation pipeline that scores and registers discovered equations
Symbolic Equation by the numbers
- 1,196 all-time installs (skills.sh)
- +33 installs in the week ending Aug 5, 2026 (Skillselion tracking)
- Ranked #930 of 16,546 AI & Agent Building skills by installs in the Skillselion catalog
- Security screen: MEDIUM risk (skills.sh audit)
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/lingzhi227/agent-research-skills --skill symbolic-equationAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 1.2k |
|---|---|
| repo stars | ★ 255 |
| Security audit | 2 / 3 scanners passed |
| Last updated | February 27, 2026 |
| Repository | lingzhi227/agent-research-skills ↗ |
How do you discover symbolic equations from scientific data?
Discover symbolic mathematical equations that model scientific systems using LLM-driven evolutionary search.
Who is it for?
ML and scientific computing developers who need interpretable symbolic formulas discovered via LLM-guided evolutionary search over equation structures.
Skip if: Production deep learning training, computer vision pipelines, or problems where black-box accuracy outweighs interpretable symbolic forms.
When should I use this skill?
The user asks to discover, evolve, or fit symbolic mathematical equations or function structures from scientific observational data.
What you get
Candidate equation functions, scored program clusters per island, and evaluated symbolic models fit to scientific input relationships.
- symbolic equation functions
- scored program candidates
By the numbers
- Implements multi-island evolutionary search with parallel LLM samplers and per-island ExperienceBuffer clusters
Files
Symbolic Equation Discovery
Discover interpretable scientific equations from data using LLM-guided evolutionary search.
Input
$0— Dataset description, variable names, and physical context
References
- LLM-SR patterns (prompts, evolution, sampling):
~/.claude/skills/symbolic-equation/references/llmsr-patterns.md
Workflow (from LLM-SR)
Step 1: Define Problem Specification
Create a specification with: 1. Input variables: Physical quantities with types (e.g., x: np.ndarray, v: np.ndarray) 2. Output variable: Target quantity to predict 3. Evaluation function: Fitness metric (typically negative MSE with parameter optimization) 4. Physical context: Domain knowledge to guide equation discovery
# Example specification
@equation.evolve
def equation(x: np.ndarray, v: np.ndarray, params: np.ndarray) -> np.ndarray:
"""Describe the acceleration of a damped nonlinear oscillator."""
return params[0] * xStep 2: Initialize Multi-Island Buffer
- Create N islands (default: 10) for population diversity
- Each island maintains independent clusters of equations
- Clusters group equations by performance signature
Step 3: Evolutionary Search Loop
Repeat until convergence or max samples: 1. Select island: Random island selection 2. Build prompt: Sample top equations from clusters (softmax-weighted by score) 3. LLM proposes: Generate new equation as improved version 4. Evaluate: Execute on test data, compute fitness score 5. Register: Add to island's cluster if valid
Step 4: Prompt Construction
Present previous equations as versioned sequence:
def equation_v0(x, v, params):
"""Initial version."""
return params[0] * x
def equation_v1(x, v, params):
"""Improved version of equation_v0."""
return params[0] * x + params[1] * v
def equation_v2(x, v, params):
"""Improved version of equation_v1."""
# LLM completes thisStep 5: Island Reset (Diversity Maintenance)
Periodically (default: every 4 hours): 1. Sort islands by best score 2. Reset bottom 50% of islands 3. Seed each reset island with best equation from a surviving island 4. Restart cluster sampling temperature
Step 6: Extract Best Equations
After search completes: 1. Collect best equation from each island 2. Rank by fitness score 3. Simplify if possible (algebraic simplification) 4. Report with physical interpretation
Cluster Sampling
Temperature-scheduled softmax over cluster scores:
temperature = T_init * (1 - (num_programs % period) / period)
probabilities = softmax(cluster_scores / temperature)- Higher temperature → more exploration
- Lower temperature → more exploitation of best clusters
- Within clusters: shorter programs are preferred (Occam's razor)
Rules
- Equations must use only standard mathematical operations
- Parameter optimization via scipy BFGS or Adam
- Fitness = negative MSE (higher is better)
- Timeout protection for equation evaluation
- No recursive equations allowed
- Physical interpretability is preferred over pure fit
Related Skills
- Upstream: data-analysis, math-reasoning
- Downstream: paper-writing-section
- See also: algorithm-design
LLM-SR Patterns
Extracted from LLM-SR codebase (llmsr/pipeline.py, sampler.py, buffer.py, evaluator.py, config.py).
LLM Instruction Prompt
You are a helpful assistant tasked with discovering mathematical function structures for scientific systems. Complete the 'equation' function below, considering the physical meaning and relationships of inputs.Multi-Island Evolutionary Algorithm
Architecture Overview
Pipeline:
├── ExperienceBuffer (multi-island)
│ ├── Island 0 → Clusters → Programs
│ ├── Island 1 → Clusters → Programs
│ ├── ...
│ └── Island N → Clusters → Programs
├── Samplers (parallel LLM callers)
│ └── get_prompt() → LLM → draw_samples()
└── Evaluators (parallel execution)
└── analyse(sample) → score → register()Main Loop (sampler.py)
def sample(self):
"""Continuously gets prompts, samples programs, sends them for analysis."""
while True:
if self._max_sample_nums and self._global_samples_nums >= self._max_sample_nums:
break
prompt = self._database.get_prompt()
samples = self._llm.draw_samples(prompt.code, self.config)
for sample in samples:
chosen_evaluator = np.random.choice(self._evaluators)
chosen_evaluator.analyse(
sample, prompt.island_id, prompt.version_generated)Prompt Construction (buffer.py)
Versioned Function Sequence
Programs from an island are formatted as an improving sequence:
def equation_v0(x, v, params):
"""Describe the acceleration of a damped oscillator."""
return params[0] * x
def equation_v1(x, v, params):
"""Improved version of equation_v0."""
return params[0] * x + params[1] * v
def equation_v2(x, v, params):
"""Improved version of equation_v1."""
return params[0] * x + params[1] * v + params[2] * x**2
def equation_v3(x, v, params):
"""Improved version of equation_v2."""
# LLM completes thisCluster-Based Selection
def get_prompt(self):
"""Constructs prompt from island clusters."""
signatures = list(self._clusters.keys())
cluster_scores = np.array(
[self._clusters[sig].score for sig in signatures])
# Temperature-scheduled softmax
period = self._cluster_sampling_temperature_period
temperature = self._cluster_sampling_temperature_init * (
1 - (self._num_programs % period) / period)
probabilities = _softmax(cluster_scores, temperature)
# Sample clusters weighted by score
functions_per_prompt = min(len(self._clusters), self._functions_per_prompt)
idx = np.random.choice(len(signatures), size=functions_per_prompt, p=probabilities)
# Sort by score ascending (worst to best)
implementations = [self._clusters[signatures[i]].sample_program() for i in idx]
indices = np.argsort([self._clusters[signatures[i]].score for i in idx])
sorted_implementations = [implementations[i] for i in indices]
return self._generate_prompt(sorted_implementations)Softmax Sampling (buffer.py)
def _softmax(logits, temperature):
"""Returns tempered softmax of 1D finite logits."""
if not np.all(np.isfinite(logits)):
raise ValueError(f'logits contains non-finite values')
result = scipy.special.softmax(logits / temperature, axis=-1)
# Fix numerical precision: ensure probabilities sum to 1
index = np.argmax(result)
result[index] = 1 - np.sum(result[0:index]) - np.sum(result[index + 1:])
return resultWithin-Cluster Sampling (Shorter Programs Preferred)
def sample_program(self):
"""Samples a program, giving higher probability to shorter programs."""
normalized_lengths = (np.array(self._lengths) - min(self._lengths)) / (
max(self._lengths) + 1e-6)
probabilities = _softmax(-normalized_lengths, temperature=1.0)
return np.random.choice(self._programs, p=probabilities)Island Reset Mechanism (buffer.py)
def reset_islands(self):
"""Resets the weaker half of islands."""
# Sort by best score (with noise to break ties)
indices_sorted = np.argsort(
self._best_score_per_island +
np.random.randn(len(self._best_score_per_island)) * 1e-6)
num_to_reset = self._config.num_islands // 2
reset_ids = indices_sorted[:num_to_reset]
keep_ids = indices_sorted[num_to_reset:]
for island_id in reset_ids:
# Create fresh island
self._islands[island_id] = Island(...)
self._best_score_per_island[island_id] = -float('inf')
# Seed with best program from a surviving island
founder_id = np.random.choice(keep_ids)
founder = self._best_program_per_island[founder_id]
self._register_program_in_island(founder, island_id, founder_scores)Reset trigger: Every reset_period seconds (default: 4 hours).
Fitness Evaluation (evaluator.py)
def analyse(self, sample, island_id, version_generated):
"""Compile and execute the hypothesis sample."""
new_function, program = _sample_to_program(
sample, version_generated, self._template, self._function_to_evolve)
scores_per_test = {}
for current_input in self._inputs:
test_output, runs_ok = self._sandbox.run(
program, self._function_to_run, self._function_to_evolve,
self._inputs, current_input, self._timeout_seconds)
if runs_ok and test_output is not None:
scores_per_test[current_input] = test_output
if scores_per_test:
self._database.register_program(new_function, island_id, scores_per_test)Score Reduction
def _reduce_score(scores_per_test):
"""Average score across all test inputs."""
return sum(scores_per_test.values()) / len(scores_per_test)Problem Specification Template
"""Specification for [PROBLEM NAME]."""
import numpy as np
from scipy.optimize import minimize
@evaluate.run
def evaluate(data: dict) -> float:
"""Evaluate equation fitness on the dataset."""
# Load data
inputs = data['inputs'] # dict of input arrays
targets = data['targets'] # target array
# Optimize parameters using BFGS
def loss(params):
pred = equation(**inputs, params=params)
return np.mean((pred - targets) ** 2)
n_params = 10 # max parameters
x0 = np.zeros(n_params)
result = minimize(loss, x0, method='L-BFGS-B')
# Return negative MSE (higher = better)
return -result.fun
@equation.evolve
def equation(x: np.ndarray, v: np.ndarray,
params: np.ndarray) -> np.ndarray:
"""Describe the acceleration of a damped nonlinear oscillator
with a driving force.
Args:
x: position (N,)
v: velocity (N,)
params: learnable parameters (K,)
Returns:
acceleration (N,)
"""
return params[0] * xConfiguration Defaults (config.py)
@dataclass
class ExperienceBufferConfig:
functions_per_prompt: int = 2 # Previous equations in prompt
num_islands: int = 10 # Population diversity
reset_period: int = 4 * 60 * 60 # 4 hours between resets
cluster_sampling_temperature_init: float = 0.1
cluster_sampling_temperature_period: int = 30_000
@dataclass
class Config:
experience_buffer: ExperienceBufferConfig
num_samplers: int = 4 # Parallel LLM callers
num_evaluators: int = 4 # Parallel evaluators
samples_per_prompt: int = 4 # Samples per LLM call
evaluate_timeout_seconds: int = 30 # Per-evaluation timeoutExample Domains
| Domain | Inputs | Output | Paper |
|---|---|---|---|
| Damped oscillator | x, v | acceleration | Physics |
| Bacterial growth | density, substrate, temp, pH | growth rate | Biology |
| Stress-strain | strain, temperature | stress | Materials |
| Oscillator + time | t, x, v | acceleration | Physics |
Related skills
How it compares
Choose symbolic-equation when interpretable symbolic formulas matter more than opaque neural network regression accuracy.
FAQ
What algorithm does symbolic-equation use?
symbolic-equation uses an LLM-SR multi-island evolutionary algorithm with parallel LLM samplers, an ExperienceBuffer clustering programs per island, and evaluators scoring proposed equation function structures.
What codebase patterns does symbolic-equation follow?
symbolic-equation follows patterns extracted from the LLM-SR codebase including pipeline.py, sampler.py, buffer.py, evaluator.py, and config.py for scientific equation discovery workflows.
Is Symbolic Equation safe to install?
skills.sh reports 2 of 3 security scanners passed. Review the Security Audits panel on this page before installing in production.