
Protein Sequence Msa
- 1.3k installs
- 2.6k repo stars
- Updated July 7, 2026
- google-deepmind/science-skills
protein-sequence-msa is a Claude Code skill that computes multiple sequence alignments on protein data using the EBI Clustal Omega service for developers building comparative protein analysis pipelines.
About
protein-sequence-msa is a science skill from google-deepmind/science-skills for computing multiple sequence alignments (MSA) on protein sequences via the EBI Clustal Omega service. The bundled Python script requires Python 3.10 or later with scienceskillscommon and python-dotenv dependencies managed through uv. Developers reach for protein-sequence-msa when agent workflows or bioinformatics pipelines need aligned protein sequence output for homology analysis, structure prediction prep, or phylogenetic comparison. The skill wraps external MSA computation so Claude, Cursor, or custom agents can submit protein FASTA inputs and retrieve alignment results without manually configuring Clustal Omega API calls.
- Runs EBI Clustal Omega MSA via REST API with automatic polling
- Accepts multi-sequence FASTA input and returns aligned output
- 15-minute polling timeout with exponential backoff
- Reuses shared HttpClient with QPS rate limiting
- Python 3.10+ script with dotenv and scienceskillscommon
Protein Sequence Msa by the numbers
- 1,268 all-time installs (skills.sh)
- +168 installs in the week ending Aug 5, 2026 (Skillselion tracking)
- Ranked #258 of 2,064 Data Science & ML skills by installs in the Skillselion catalog
- Security screen: LOW risk (skills.sh audit)
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/google-deepmind/science-skills --skill protein-sequence-msaAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 1.3k |
|---|---|
| repo stars | ★ 2.6k |
| Security audit | 3 / 3 scanners passed |
| Last updated | July 7, 2026 |
| Repository | google-deepmind/science-skills ↗ |
How do you compute protein multiple sequence alignments?
Compute multiple sequence alignments on protein data using the EBI Clustal Omega service.
Who is it for?
Bioinformatics developers building protein analysis pipelines who need automated MSA computation via EBI Clustal Omega from agent or Python workflows.
Skip if: DNA or RNA-only alignment tasks, manual alignment editing in desktop tools, or workflows requiring local HMMER execution without external API calls.
When should I use this skill?
A developer asks to compute protein multiple sequence alignments, run Clustal Omega on protein FASTA input, or integrate MSA into a bioinformatics pipeline.
What you get
Multiple sequence alignment output from EBI Clustal Omega for submitted protein FASTA sequences.
- Multiple sequence alignment output
- Aligned protein sequence file
By the numbers
- Requires Python >=3.10 with scienceskillscommon and python-dotenv dependencies
Files
Prerequisites
1. `uv`: Read the uv skill and follow its Setup instructions to ensure uv is installed and on PATH. 2. User Notification: If LICENSE_NOTIFICATION.txt does not already exist in this skill directory then (1) prominently notify the user to check the terms at https://www.ebi.ac.uk/jdispatcher/msa/clustalo and https://www.ebi.ac.uk/about/terms-of-use/, then (2) create the file recording the notification text and timestamp. 3. `.env` file: Make sure the .env file exists in your home directory. Create one if it does not exist. 4. `USER_EMAIL` (optional but recommended): Recommended by the EBI for Clustal Omega job tracking, but the skill works without it. If the variable is missing from .env, do NOT ask the user to paste it into the chat (this would leak the value into the agent's context). Instead, give the user this command — substituting `ENV_FILE` with the resolved literal path to the `.env` file:
printf "Enter contact email: " && read email && echo "USER_EMAIL=$email" >> "ENV_FILE" && echo "Saved."The scripts load credentials automatically via dotenv. NEVER read, print, or inspect the .env file or its variables (e.g. no cat, grep, echo, printenv, or os.environ.get on keys). Credentials must stay out of the agent's context.
Core Rules
- Use the Wrapper: ALWAYS execute the alignment using
scripts/msa_align.py rather than writing your own curl or custom Python requests. The script automatically enforces the required rate limit to respect EBI's Terms of Use.
- Notification: If this skill is used, ensure this is mentioned in the
output.
- Always state the method: Every report must clearly state that the
alignment was performed using EBI Clustal Omega.
- No Hallucinations: Do NOT invent alignments or conservation metrics.
Report only what is present in the alignment file.
Goal
Take a file containing multiple protein sequences in FASTA format, perform multiple sequence alignment using the EBI Clustal Omega API, save the resulting alignment locally for future programmatic analysis, and interpret the results towards addressing the user's specific research objective (e.g., assessing similarity, identifying conserved domains, or analyzing key residues).
Instructions
1. Prepare Input File: The input must be a plain text file containing two or more protein sequences in FASTA format. Each sequence header must start with a > symbol. Example:
>Sequence_1_Name
MQIFVKTLTGKTITLEVEPSDTIENVKAKIQDKEGIPPDQ
QRLIFAGKQLEDGRTLSDYNIQKESTLHLVLRLRGG
>Sequence_2_Name
MQIFVKTLTGKTITLEVEPSDTIENVKAKIQDKEGIPPDQ
QRLIFAGKQLEDGRTLSDYNIQKESTLHLVLRLRGG2. Execute Alignment: Run the alignment script:
uv run scripts/msa_align.py <INPUT_FASTA> -o <OUTPUT_FILE>Always specify the output file with -o or --output.
3. Interpret and Report Results: Analyze the Clustal Omega alignment by selecting metrics and mapping strategies aligned with the research objective. Note that while Clustal Omega produces a Global Alignment, pairwise metrics can be extracted to evaluate specific relationships within the set:
- Identity Metric Options: The choice of denominator determines how
insertions/deletions (gaps) affect the final percentage. Select the most appropriate calculation based on the biological context:
- Pairwise - Sequence Coverage: `(Identical Residue Matches) /
(Length of Shorter Sequence)`. Use when determining if a specific domain or fragment is fully preserved within a larger protein. This ignores gaps in the longer sequence, focusing purely on the "content" of the shorter one.
- Pairwise - Global Identity: `(Identical Residue Matches) /
(Total Alignment Columns)`. Use when comparing full-length sequences of similar expected length. This is the most conservative metric; it penalizes for all gaps (indels) introduced by any sequence in the MSA.
- Pairwise - Overlap Identity: `(Identical Residue Matches) /
(Total Alignment Columns - Terminal Gaps)`. Use when comparing a fragment to a full-length protein or when sequences have long unaligned "tails." This focuses on similarity only where the sequences physically overlap.
- Multisequence - Conservation Index: `(Fully Conserved Columns) /
(Total Alignment Columns)`. Use for quantifying the percentage of residues that are 100% identical across the entire alignment set. This identifies the core evolutionary signature of the protein family.
- Feature Mapping: Leverage known biological data from specific
sequences to ground the analysis:
- Knowledge Gathering: Identify relevant known sites or regions
(e.g., catalytic residues, binding motifs) from your input or via external tools.
- Coordinate Projection: Map these features onto the corresponding
Column Indices of the alignment.
- Targeted Discussion: Use these columns to drive the assessment:
- Local Conservation: Analyze if the known functional residues
are invariant across the set.
- Region-Specific Metrics: Calculate identity/similarity
specifically within the mapped functional regions rather than the whole sequence.
- Goal Contribution: Discuss how this data contributes to your
goal, e.g. using conservation to corroborate a prediction or divergence to reject a functional hypothesis.
References
- Multiple Sequence Alignment: https://www.ebi.ac.uk/jdispatcher/msa/clustalo
# Copyright 2026 Google LLC
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
# /// script
# requires-python = ">=3.10"
# dependencies = [
# "scienceskillscommon",
# "python-dotenv",
# ]
# [tool.uv.sources]
# scienceskillscommon = { path = "../../scienceskillscommon" }
# ///
"""Runs EBI Clustal Omega for MSA computation.
Takes a file with multiple sequences and provides the alignment.
"""
import argparse
import os
import sys
import time
import urllib.parse
import dotenv
from science_skills.skills.scienceskillscommon import http_client
_POLLING_TIMEOUT_SECS = 15 * 60 # 15 minutes.
_CLIENT = http_client.HttpClient(
"https://www.ebi.ac.uk/Tools/services/rest/clustalo/", qps=1
)
def _prepare_payload(email: str, title: str, sequences: str) -> bytes:
"""Prepares the payload for the EBI Clustal Omega API."""
params = {
"email": email,
"title": title,
"sequence": sequences,
}
return urllib.parse.urlencode(params).encode("utf-8")
def _align_sequences(
*, input_file: str, output_file: str, dry_run: bool = False
) -> None:
"""Runs EBI Clustal Omega alignment for sequences in a FASTA file.
This function takes a FASTA formatted file, submits the sequences to the
EBI Clustal Omega web service, polls for the alignment completion, and
saves the resulting alignment in FASTA format to the specified output file.
Args:
input_file: Path to the input file containing sequences in FASTA format.
output_file: Path where the resulting MSA in FASTA format will be saved.
dry_run: If True, print the payload and exit without submitting the job.
"""
if not os.path.exists(input_file):
print(f"[!] Error: Input file not found: {input_file}")
sys.exit(1)
max_size_bytes = 4 * 1024 * 1024 # 4 MB
file_size = os.path.getsize(input_file)
if file_size > max_size_bytes:
print(
"[!] Error: At most 4 MB file size supported. Found"
f" {file_size / (1024 * 1024):.2f} MB."
)
sys.exit(1)
with open(input_file, "r") as f:
sequences = f.read().strip()
if not sequences:
print("[!] Error: Empty input file.")
sys.exit(1)
num_sequences = sequences.count(">")
if num_sequences < 2:
print(f"[!] Error: At least 2 sequences required. Found {num_sequences}.")
sys.exit(1)
if num_sequences > 4000:
print(
f"[!] Error: At most 4000 sequences supported. Found {num_sequences}."
)
sys.exit(1)
print("[*] Submitting sequences to EBI Clustal Omega API...")
# 1. Submit Job
user_email = os.environ.get("USER_EMAIL")
if not user_email:
print("[!] Error: USER_EMAIL environment variable is required.")
sys.exit(1)
data = _prepare_payload(user_email, "MSA", sequences)
if dry_run:
print(data)
sys.exit(0)
job_id = _CLIENT.fetch_text(
"run", method="POST", data=data, headers={"Accept": "text/plain"}
).strip()
print(f"[*] Job ID generated: {job_id}")
# 2. Poll the server
print("[*] Polling server for completion...")
start_time = time.time()
while time.time() - start_time < _POLLING_TIMEOUT_SECS:
status = _CLIENT.fetch_text(
f"status/{job_id}", headers={"Accept": "text/plain"}, timeout=20
).strip()
sys.stdout.write(".")
sys.stdout.flush()
if status == "FINISHED":
print("\n[*] Job marked as FINISHED.")
break
elif status in ["ERROR", "FAILURE", "NOT_FOUND"]:
print(f"\n[!] Job failed with status: {status}")
sys.exit(1)
time.sleep(10)
else:
print(f"\n[!] Job timed out after {_POLLING_TIMEOUT_SECS // 60} minutes.")
sys.exit(1)
# 3. Fetch Results
print("\n[*] Job complete. Fetching results...\n")
result_text = _CLIENT.fetch_text(f"result/{job_id}/fa", timeout=60)
with open(output_file, "w") as f:
f.write(result_text)
print(f"[*] Alignment results saved to: {output_file}")
def main() -> None:
dotenv.load_dotenv(os.path.expanduser("~/.env"))
parser = argparse.ArgumentParser(
description="MSA computation using EBI Clustal Omega."
)
parser.add_argument(
"input", help="Path to FASTA file containing multiple sequences"
)
parser.add_argument(
"-o",
"--output",
required=True,
help="Path to save the output alignment file",
)
parser.add_argument(
"--dry-run",
action="store_true",
help="Dry run: print payload and exit without submitting job",
)
args = parser.parse_args()
_align_sequences(
input_file=args.input, output_file=args.output, dry_run=args.dry_run
)
if __name__ == "__main__":
main()
Related skills
FAQ
What service does protein-sequence-msa use?
protein-sequence-msa computes multiple sequence alignments using the EBI Clustal Omega service. Developers submit protein FASTA sequences and receive alignment output suitable for downstream homology or structure analysis.
What Python version does protein-sequence-msa require?
protein-sequence-msa requires Python 3.10 or later with scienceskillscommon and python-dotenv dependencies. The script is managed through uv in the google-deepmind/science-skills repository.
Is Protein Sequence Msa safe to install?
skills.sh reports 3 of 3 security scanners passed. Review the Security Audits panel on this page before installing in production.