
Gtex Database
- 1.2k installs
- 2.6k repo stars
- Updated July 7, 2026
- google-deepmind/science-skills
gtex-database is a Google DeepMind science skill that queries GTEx Portal API V2 for median gene expression and eQTL data across 54 tissues via a rate-limited CLI wrapper for genomics developers.
About
gtex-database is a Google DeepMind science-skills module that wraps GTEx Portal API V2 for baseline RNA expression and expression quantitative trait loci. The gtex_cli.py script exposes subcommands to resolve gene symbols to versioned GENCODE IDs, fetch median TPM expression, list top-expressed tissues, retrieve gene-level eQTLs, and search eQTLs in chromosomal regions—with pagination capped at 250 items per page and sequential fetching per GTEx Terms of Use. First use requires license notification to .licenses/gtex_database_LICENSE.txt. Outputs are JSON files under /tmp/ for agent consumption. Reach for gtex-database when you need normal-tissue expression atlases or variant-gene regulatory context—not protein-level expression, diseased tissue samples, or fetal expression data.
- CLI wrapper for GTEx API V2 with built-in rate limiting and pagination
- Respects GTEx Portal Terms of Use through sequential fetching
- Reuses shared HttpClient from scienceskillscommon library
- Supports tissue caching to minimize repeated network calls
- Outputs clean JSON for downstream analysis or agent consumption
Gtex Database by the numbers
- 1,244 all-time installs (skills.sh)
- +167 installs in the week ending Aug 5, 2026 (Skillselion tracking)
- Ranked #269 of 2,064 Data Science & ML skills by installs in the Skillselion catalog
- Security screen: LOW risk (skills.sh audit)
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/google-deepmind/science-skills --skill gtex-databaseAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 1.2k |
|---|---|
| repo stars | ★ 2.6k |
| Security audit | 3 / 3 scanners passed |
| Last updated | July 7, 2026 |
| Repository | google-deepmind/science-skills ↗ |
How do you query GTEx gene expression by tissue?
Query the GTEx Portal V2 API for human gene expression and tissue data directly from the terminal or inside an agent workflow.
Who is it for?
Genomics developers who need GTEx baseline expression and eQTL context through a scripted, terms-compliant API V2 wrapper.
Skip if: Skip gtex-database for tumor expression, protein-level measurements, or embryonic/fetal transcriptomic data outside GTEx's adult normal-tissue atlas.
When should I use this skill?
User needs GTEx median TPM expression, tissue-specific expression ranking, or eQTL data for a gene or genomic region.
What you get
JSON files with GENCODE IDs, median TPM expression, top tissues, and significant eQTL records
- GENCODE ID JSON
- median expression JSON
- eQTL result JSON
By the numbers
- Queries expression data across 54 non-diseased GTEx tissue sites
- Limits GTEx API pagination to 250 items per page
Files
GTEx Database Integration
This skill retrieves transcriptomics data (RNA expression baselines) and expression Quantitative Trait Loci (eQTLs) from the GTEx Portal API V2. It provides access to median TPM (Transcripts Per Million) values for genes and significant eQTLs for variants across 54 human tissue sites.
Prerequisites
1. `uv`: Read the uv skill and follow its Setup instructions to ensure uv is installed and on PATH. 2. User Notification: If LICENSE_NOTIFICATION.txt does not already exist in this skill directory then (1) prominently notify the user to check the terms at https://gtexportal.org/home/license and https://gtexportal.org/home/documentationPage#gtexApi, then (2) create the file recording the notification text and timestamp.
When to Use
Use this skill when you need to:
- Map a gene symbol to its Versioned GENCODE ID.
- Retrieve the baseline median expression level (in TPM) of a gene across
various tissues.
- Find the top tissues where a particular gene is most highly expressed.
- Fetch significant single-tissue eQTLs for a variant or within a chromosomal
window.
- Get all significant eQTLs associated with a specific gene.
- Contextualise a variant within GWAS loci using eQTL data.
Do NOT use when you need to:
- Query for protein-level expression or post-translational modifications
(PTMs). GTEx only measures mRNA abundance.
- Query gene expression in diseased tissues (e.g., tumor samples, cirrhosis).
GTEx is a baseline atlas of normal, non-diseased tissues.
- Query embryonic or fetal gene expression. GTEx donors are adults only.
Core Rules
CRITICAL: You MUST respect GTEx Portal API Terms of Use.
- Use the Wrapper: ALWAYS execute the provided helper scripts to query the
database rather than accessing the database directly. The scripts automatically enforce the required rate limit gracefully.
- Limit requests to maximum 250 items per page where applicable.
- Notification: If this skill is used, ensure this is mentioned in the
output.
Command Selection Guide
Pick the right command on the first try. Match the user's input to the correct subcommand below.
- Map a gene symbol to GENCODE ID:
resolve-gencode-id - Get median expression (TPM) for a gene:
get-median-expression - Find tissues with highest expression for a gene:
get-top-expressed-tissues - Get all eQTLs for a specific gene:
get-gene-eqtls - Find eQTLs within a chromosomal region:
get-eqtls-in-region
Quick Start
# Map the TNF gene symbol to its GENCODE ID
uv run scripts/gtex_cli.py resolve-gencode-id TNF --output /tmp/tnf_id.json
# Get median expression of a gene by GENCODE ID
uv run scripts/gtex_cli.py get-median-expression ENSG00000232810.2 --output /tmp/tnf_expr.jsonAll subcommands write JSON to disk. Always save output in the /tmp/ directory. The default output file is /tmp/gtex_output.json if --output is not specified.
Commands
1. resolve-gencode-id — Gene Symbol → GENCODE ID
Maps a standard gene symbol (e.g., "JUN", "TNF") to its Versioned GENCODE ID. This ID is required for all other expression and eQTL calls.
uv run scripts/gtex_cli.py resolve-gencode-id TNF --output /tmp/tnf_id.jsonArguments:
-
gene_symbol(positional): The standard gene symbol (e.g., "TNF"). -
--output: Output file path (default:/tmp/gtex_output.json).
2. get-median-expression — Get Median Expression (TPM)
Retrieves the median TPM for a gene across all 54 GTEx tissue sites or specified tissues.
uv run scripts/gtex_cli.py get-median-expression ENSG00000232810.2 \
--tissues "Whole Blood,Spleen" --output /tmp/expr.jsonArguments:
-
gencode_id(positional): The Versioned GENCODE ID. -
--tissues: Comma-separated list of tissue IDs (optional, defaults to all
54 tissues).
-
--output: Output file path (default:/tmp/gtex_output.json).
3. get-top-expressed-tissues — Get Top Expressed Tissues
Returns the n tissues with the highest median expression for the target gene.
uv run scripts/gtex_cli.py get-top-expressed-tissues ENSG00000232810.2 \
--n 5 --output /tmp/top_tissues.jsonArguments:
-
gencode_id(positional): The Versioned GENCODE ID. -
--n: Number of top tissues to return (default: 5). -
--output: Output file path.
4. get-gene-eqtls — Get All eQTLs for a Gene
Returns every significant eQTL associated with the gene across specified tissues.
uv run scripts/gtex_cli.py get-gene-eqtls ENSG00000232810.2 \
--tissues "Whole Blood" --output /tmp/eqtls.jsonArguments:
-
gencode_id(positional): The Versioned GENCODE ID. -
--tissues: Comma-separated list of tissue IDs (optional, defaults to all). -
--output: Output file path.
5. get-eqtls-in-region — Get eQTLs in Chromosomal Region
Returns all significant single-tissue eQTLs within a chromosomal window (up to 8Mb).
uv run scripts/gtex_cli.py get-eqtls-in-region chr17 7000000 7100000 "Esophagus - Muscularis" \
--output /tmp/region_eqtls.jsonArguments:
-
chromosome(positional): Chromosome name (e.g.,chr17). -
start(positional): Start position. -
end(positional): End position (max 8Mb from start). -
tissue_id(positional): The target tissue ID. -
--output: Output file path.
Typical Workflows
Identify highest expressing tissues for a gene
# Step 1: Map symbol to GENCODE ID
uv run scripts/gtex_cli.py resolve-gencode-id GATA4 --output /tmp/gata4_id.json
# Step 2: Query for top tissues using the resolved ID
uv run scripts/gtex_cli.py get-top-expressed-tissues <gencode_id> --n 5 \
--output /tmp/gata4_top.json# Copyright 2026 Google LLC
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
"""CLI wrapper for GTEx API V2.
Follows GTEx Portal Terms of Use by fetching sequentially and handling
pagination.
"""
# /// script
# requires-python = ">=3.10"
# dependencies = [
# "scienceskillscommon",
# ]
# [tool.uv.sources]
# scienceskillscommon = { path = "../../scienceskillscommon" }
# ///
import argparse
import json
import sys
import urllib.parse
from science_skills.skills.scienceskillscommon import http_client
BASE_URL = 'https://gtexportal.org/api/v2'
DATASET_ID = 'gtex_v10'
GENCODE_VERSION = 'v39'
CLIENT = http_client.HttpClient(BASE_URL, qps=1.0)
# Optional cache for tissues to avoid fetching repeatedly
TISSUE_CACHE = None
def _fetch(url, params=None):
"""Fetches URL, with optional query parameters, using HttpClient."""
if params:
# Filter out None values
params = {k: v for k, v in params.items() if v is not None}
query_string = urllib.parse.urlencode(params, doseq=True)
full_url = f'{url}?{query_string}'
else:
full_url = url
return CLIENT.fetch_json(full_url)
def _fetch_paginated(url, params=None):
"""Fetches all pages from a paginated endpoint."""
if params is None:
params = {}
all_data = []
page = 0
while True:
params['page'] = page
response = _fetch(url, params)
# Some endpoints don't use 'data' wrapping or 'paging_info'
if 'paging_info' not in response:
return response
data = response.get('data', [])
all_data.extend(data)
paging = response.get('paging_info', {})
total_pages = paging.get('numberOfPages', 0)
if page >= total_pages - 1:
break
page += 1
return all_data
def get_tissue_mapping():
"""Fetches the list of tissues and maps names to tissueSiteDetailId."""
global TISSUE_CACHE
if TISSUE_CACHE is not None:
return TISSUE_CACHE
url = f'{BASE_URL}/dataset/tissueSiteDetail'
data = _fetch_paginated(url, {'datasetId': DATASET_ID})
mapping = {}
for t in data:
id_ = t.get('tissueSiteDetailId')
name = t.get('tissueSiteDetail')
if id_ and name:
mapping[id_.lower()] = id_
mapping[name.lower()] = id_
# Allow "Esophagus - Muscularis" instead of "Esophagus_Muscularis", etc.
mapping[id_.replace('_', ' ').lower()] = id_
mapping[name.replace('-', ' ').lower()] = id_
TISSUE_CACHE = mapping
return mapping
def resolve_tissue(tissue_str):
mapping = get_tissue_mapping()
cleaned = tissue_str.strip().lower()
if cleaned in mapping:
return mapping[cleaned]
# Try more fuzzy matching if needed
cleaned_no_hyphen = cleaned.replace('-', ' ')
if cleaned_no_hyphen in mapping:
return mapping[cleaned_no_hyphen]
sys.stderr.write(f"Error: Unknown tissue '{tissue_str}'.\n")
sys.exit(1)
def resolve_gencode_id(gene_symbol, output_file):
"""Maps a standard gene symbol to its Versioned GENCODE ID."""
url = f'{BASE_URL}/reference/gene'
data = _fetch_paginated(
url, {'geneId': gene_symbol, 'gencodeVersion': GENCODE_VERSION}
)
if not data:
sys.stderr.write(f"Error: Could not find GENCODE ID for '{gene_symbol}'.\n")
sys.exit(1)
# Return the first matching exact symbol if possible, else just the first one
best_match = data[0]
for d in data:
if d.get('geneSymbol', '').lower() == gene_symbol.lower():
best_match = d
break
result = {
'gene_symbol': best_match.get('geneSymbol'),
'gencode_id': best_match.get('gencodeId'),
'chromosome': best_match.get('chromosome'),
'start': best_match.get('start'),
'end': best_match.get('end'),
'gene_type': best_match.get('geneType'),
}
with open(output_file, 'w') as f:
json.dump(result, f, indent=2)
def get_median_expression(gencode_id, tissues, output_file):
"""Returns the median expression (TPM) for a gene across tissues."""
url = f'{BASE_URL}/expression/medianGeneExpression'
params = {'gencodeId': gencode_id, 'datasetId': DATASET_ID}
if tissues:
tissue_list = [t.strip() for t in tissues.split(',')]
resolved_tissues = [resolve_tissue(t) for t in tissue_list]
# The API accepts multiple tissueSiteDetailId parameters.
# urllib.urlencode with doseq=True handles list correctly.
params['tissueSiteDetailId'] = resolved_tissues
data = _fetch_paginated(url, params)
with open(output_file, 'w') as f:
json.dump(data, f, indent=2)
def get_top_expressed_tissues(gencode_id, n, output_file):
url = f'{BASE_URL}/expression/medianGeneExpression'
params = {'gencodeId': gencode_id, 'datasetId': DATASET_ID}
data = _fetch_paginated(url, params)
# Sort by median expression, descending
sorted_data = sorted(data, key=lambda x: x.get('median', 0), reverse=True)
top_n = sorted_data[:n]
with open(output_file, 'w') as f:
json.dump(top_n, f, indent=2)
def get_gene_eqtls(gencode_id, tissues, output_file):
"""Returns all significant eQTLs for a gene across tissues."""
url = f'{BASE_URL}/association/singleTissueEqtl'
params = {'gencodeId': gencode_id, 'datasetId': DATASET_ID}
if tissues:
tissue_list = [t.strip() for t in tissues.split(',')]
resolved_tissues = [resolve_tissue(t) for t in tissue_list]
params['tissueSiteDetailId'] = resolved_tissues
data = _fetch_paginated(url, params)
with open(output_file, 'w') as f:
json.dump(data, f, indent=2)
def get_eqtls_in_region(chromosome, start, end, tissue, output_file):
"""Returns all significant eQTLs for a region in a tissue."""
url = f'{BASE_URL}/association/singleTissueEqtlByLocation'
resolved_tissue = resolve_tissue(tissue)
params = {
'chromosome': chromosome,
'start': start,
'end': end,
'tissueSiteDetailId': resolved_tissue,
'datasetId': DATASET_ID,
}
# This endpoint is NOT paginated per API docs and does not return paging_info.
data = _fetch(url, params)
with open(output_file, 'w') as f:
json.dump(data, f, indent=2)
def main():
parser = argparse.ArgumentParser(description='GTEx Portal API V2 CLI')
subparsers = parser.add_subparsers(dest='command', required=True)
p_resolve = subparsers.add_parser(
'resolve-gencode-id',
help='Map a standard gene symbol to its Versioned GENCODE ID',
)
p_resolve.add_argument('gene_symbol', help='Gene symbol (e.g. TNF)')
p_resolve.add_argument('--output', default='/tmp/gtex_output.json')
p_median = subparsers.add_parser(
'get-median-expression', help='Get median expression (TPM)'
)
p_median.add_argument('gencode_id', help='Versioned GENCODE ID')
p_median.add_argument('--tissues', help='Comma-separated list of tissue IDs')
p_median.add_argument('--output', default='/tmp/gtex_output.json')
p_top = subparsers.add_parser(
'get-top-expressed-tissues', help='Get top expressed tissues for a gene'
)
p_top.add_argument('gencode_id', help='Versioned GENCODE ID')
p_top.add_argument(
'--n', type=int, default=5, help='Number of tissues to return'
)
p_top.add_argument('--output', default='/tmp/gtex_output.json')
p_eqtls = subparsers.add_parser(
'get-gene-eqtls', help='Get all significant eQTLs for a gene'
)
p_eqtls.add_argument('gencode_id', help='Versioned GENCODE ID')
p_eqtls.add_argument('--tissues', help='Comma-separated list of tissue IDs')
p_eqtls.add_argument('--output', default='/tmp/gtex_output.json')
p_region = subparsers.add_parser(
'get-eqtls-in-region', help='Get significant eQTLs in a region'
)
p_region.add_argument('chromosome', help='Chromosome (e.g. chr17)')
p_region.add_argument('start', type=int, help='Start position')
p_region.add_argument('end', type=int, help='End position')
p_region.add_argument('tissue_id', help='Target tissue ID')
p_region.add_argument('--output', default='/tmp/gtex_output.json')
args = parser.parse_args()
if args.command == 'resolve-gencode-id':
resolve_gencode_id(args.gene_symbol, args.output)
elif args.command == 'get-median-expression':
get_median_expression(args.gencode_id, args.tissues, args.output)
elif args.command == 'get-top-expressed-tissues':
get_top_expressed_tissues(args.gencode_id, args.n, args.output)
elif args.command == 'get-gene-eqtls':
get_gene_eqtls(args.gencode_id, args.tissues, args.output)
elif args.command == 'get-eqtls-in-region':
get_eqtls_in_region(
args.chromosome, args.start, args.end, args.tissue_id, args.output
)
else:
parser.print_help()
if __name__ == '__main__':
main()
Related skills
How it compares
Use gtex-database over manual API calls when you need terms-compliant, rate-limited GTEx expression and eQTL JSON for agent workflows.
FAQ
How many tissues does gtex-database cover?
gtex-database retrieves baseline median RNA expression and eQTL data across 54 non-diseased human tissue sites via GTEx Portal API V2. GTEx measures mRNA abundance only, not protein-level expression.
Which gtex-database commands should you run first?
gtex-database starts with resolve-gencode-id to map a gene symbol to a versioned GENCODE ID. Subsequent calls use get-median-expression, get-top-expressed-tissues, get-gene-eqtls, or get-eqtls-in-region depending on the query.
Is Gtex Database safe to install?
skills.sh reports 3 of 3 security scanners passed. Review the Security Audits panel on this page before installing in production.