Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
neo4j-contrib avatar

Neo4j Gds Skill

  • 391 installs
  • 101 repo stars
  • Updated August 3, 2026
  • neo4j-contrib/neo4j-skills

neo4j-gds-skill is a Claude Code skill that runs Neo4j Graph Data Science projections and core algorithms from Python or Cypher with memory estimation and embedding write-back patterns for developers building graph analy

About

neo4j-gds-skill is an agent skill from neo4j-contrib/neo4j-skills for Neo4j Graph Data Science on Aura Pro, self-managed, local, or offline DBMS instances with the GDS plugin installed. The skill covers native and Cypher graph projection, execution modes—stream, stats, mutate, and write—and seven core algorithms: PageRank, Louvain, WCC, Betweenness Centrality, Node Similarity, FastRP, and KNN. It documents the FastRP-to-KNN recommendation pipeline, writing node embeddings for Neo4j vector indexes, and memory estimation before large projections via the graphdatascience Python client. Developers reach for neo4j-gds-skill when shipping community detection, centrality scoring, or structural similarity search without guessing GDS mode semantics or blowing heap on unestimated projections.

  • Native and Cypher graph projection with stream/stats/mutate/write execution modes
  • Core algorithms: PageRank, Louvain, WCC, Betweenness, Node Similarity, FastRP, KNN
  • FastRP → KNN recommendation pipeline with embeddings written for vector index follow-up
  • Memory estimation before large projections and catalog ops: project, list, drop, subgraph filter
  • GDS Python client v2 with v1 fallback; documents OOM and licensing error mitigations

Neo4j Gds Skill by the numbers

  • 391 all-time installs (skills.sh)
  • +30 installs in the week ending Aug 4, 2026 (Skillselion tracking)
  • Ranked #151 of 911 Databases skills by installs in the Skillselion catalog
  • Security screen: LOW risk (skills.sh audit)
  • Data as of Aug 4, 2026 (Skillselion catalog sync)
npx skills add https://github.com/neo4j-contrib/neo4j-skills --skill neo4j-gds-skill

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs391
repo stars101
Security audit3 / 3 scanners passed
Last updatedAugust 3, 2026
Repositoryneo4j-contrib/neo4j-skills

How do you run Neo4j GDS algorithms from Python?

Run Neo4j Graph Data Science projections and core algorithms from Python or Cypher with memory estimation and embedding write-back patterns.

Who is it for?

Backend engineers adding graph analytics, community detection, or embedding pipelines on Neo4j with the GDS plugin installed.

Skip if: Teams running vanilla Cypher CRUD without GDS installed or needing OLTP-only graph queries without algorithm projections.

When should I use this skill?

A task involves GDS projection, PageRank, Louvain, FastRP, KNN pipelines, or memory estimation for Neo4j graph algorithms.

What you get

GDS graph projections, algorithm results in stream/mutate/write modes, and node embeddings written for vector indexes.

  • GDS projections
  • Algorithm result sets
  • Node embeddings for vector indexes

By the numbers

  • Covers 7 core GDS algorithms: PageRank, Louvain, WCC, Betweenness Centrality, Node Similarity, FastRP, and KNN
  • Documents 4 GDS execution modes: stream, stats, mutate, and write

Files

SKILL.mdMarkdownGitHub ↗

When to Use

  • Running GDS algorithms against embedded GDS plugin through Python client (graphdatascience)
  • Running GDS algorithms through CALL gds.* Cypher procedures
  • Aura Pro, self-managed Neo4j, local Neo4j, or offline DBMS with GDS plugin installed
  • Projecting named in-memory graphs, running centrality/community/similarity/path/embedding algorithms
  • Chaining algorithms via mutate mode; building FastRP → KNN pipelines
  • Writing node embeddings for Neo4j vector indexes / structural similarity search
  • Memory estimation before large graph operations

When NOT to Use

  • Aura Graph Analytics Sessions / AGA / `GdsSessions` / `AuraGraphDataScience`neo4j-aura-graph-analytics-skill
  • AuraDB Cypher API with `{ memory: ... }` or `{ sessionId: ... }`neo4j-aura-graph-analytics-skill
  • Cypher query authoringneo4j-cypher-skill
  • Driver/connection setupneo4j-driver-python-skill
  • GraphRAG retrievalneo4j-graphrag-skill
  • Creating/querying vector indexes over written embeddingsneo4j-vector-index-skill
ContextUse
Aura Pro with GDS pluginThis skill
Self-managed/local/offline Neo4j with GDS pluginThis skill
AuraDB serverless analytics sessionneo4j-aura-graph-analytics-skill
Self-managed Neo4j attached to AGA sessionneo4j-aura-graph-analytics-skill
Non-Neo4j data sourceneo4j-aura-graph-analytics-skill

---

Pre-flight

Use only with embedded GDS plugin.

from graphdatascience import GraphDataScience

gds = GraphDataScience("neo4j+s://xxx.databases.neo4j.io", auth=("neo4j", "pw"), aura_ds=True)
gds = GraphDataScience("bolt://localhost:7687", auth=("neo4j", "password"))
print(gds.server_version())
RETURN gds.version() AS gds_version

If Unknown function 'gds.version' → GDS plugin unavailable. AuraDB serverless analytics → neo4j-aura-graph-analytics-skill. Self-managed/local → install or enable GDS plugin.

pip install graphdatascience              # Python client
pip install graphdatascience[rust_ext]    # 3–10× faster serialization

Compatibility: graphdatascience v1.22 — GDS >= 2.6 and < 2.28 / < 2026.6, Python >= 3.10 and < 3.15, Neo4j Driver >= 4.4.12 and < 7.0.

V2 rules:

  • Prefer gds.v2.* when endpoint exists.
  • Use snake_case endpoints and parameters: page_rank, fast_rp, mutate_property, write_property.
  • Use typed result attributes: result.write_millis, not result["writeMillis"].
  • Use v1 if v2 endpoint missing/incompatible; label fallback.

---

Graph Catalog Operations

Native Projection

CALL gds.graph.project(
  'myGraph',
  ['Person', 'City'],
  { KNOWS: { orientation: 'UNDIRECTED' }, LIVES_IN: {} }
)
YIELD graphName, nodeCount, relationshipCount
G, result = gds.v2.graph.project("myGraph", "Person", "KNOWS")
print(result.node_count, result.relationship_count)

G, result = gds.v2.graph.project(
    "myGraph",
    {"Person": {"properties": ["age", "score"]}, "City": {}},
    {"KNOWS": {"orientation": "UNDIRECTED"}, "LIVES_IN": {"properties": ["since"]}}
)

Native projection: plugin/simple Python-client workflow only. AGA Sessions → neo4j-aura-graph-analytics-skill. V1 fallback: gds.graph.project(...).

Cypher Projection (use for new Cypher workflows, filters, transforms)

G, result = gds.graph.cypher.project(
    """
    MATCH (source:Person)-[r:KNOWS]->(target:Person)
    WHERE source.active = true
    RETURN gds.graph.project($graph_name, source, target,
        { sourceNodeProperties: source { .score }, relationshipType: 'KNOWS' })
    """,
    database="neo4j", graph_name="activeGraph"
)

gds.graph.cypher.project must end with one RETURN gds.graph.project(...) clause. If validation fails: use gds.run_cypher(...), then gds.graph.get("graphName"). Use v1 gds.graph.cypher.project(...) if v2 graph projection cannot express required filter/transform.

AGA Sessions → neo4j-aura-graph-analytics-skill; never use plugin Cypher projection.

Undirected Projection

Native projection: set orientation: 'UNDIRECTED' per relationship type. Plugin Cypher projection: set undirectedRelationshipTypes: ['*'] in fifth gds.graph.project(...) config argument.

Leiden is defined for directed and undirected graphs. Project undirected relationships when community structure is naturally symmetric.

Inspect and Drop

G.node_count()              # 12_043
G.relationship_count()      # 87_211
G.node_properties()         # projected + mutated properties by label
G.relationship_properties() # projected + mutated properties by type
G.size_in_bytes()
gds.v2.graph.drop(G)        # frees JVM heap

G = gds.v2.graph.get("myGraph")       # re-attach to existing projection

gds.v2.graph.list()

Memory Estimation — run before large projections and algorithms

CALL gds.graph.project.estimate(['Person'], 'KNOWS')
YIELD requiredMemory, bytesMin, bytesMax, nodeCount, relationshipCount
G, project_result = gds.v2.graph.project("myGraph", "Person", "KNOWS")
print(project_result.node_count)

# Algorithm estimation:
est = gds.v2.page_rank.estimate(G, damping_factor=0.85)
print(est.required_memory)

Projection estimate fallback: use v1 gds.graph.project.estimate(...) if v2 estimate endpoint unavailable.

---

Execution Modes

ModeSide effectReturnsUse when
streamNoneRow per node/pairInspect results; top-N
statsNoneSingle aggregate rowSummary/convergence check
mutateAdds node property or relationship type/property to in-memory graph onlyStats rowChain algorithms
writePersists node property or relationship to Neo4j DBStats rowFinal step — make queryable

Pattern: stream to verify → mutate to chain → write to persist.

mutate_property must not exist in the in-memory graph. Relationship algorithms such as KNN also require mutate_relationship_type. After write, re-project to use written properties in subsequent GDS calls (in-memory graph does not see DB writes).

---

gds.util.asNode() — Enrich Stream Results

stream mode yields nodeId (internal GDS integer). gds.util.asNode(nodeId) translates it back to the DB node so you can access properties.

// Single property
CALL gds.pageRank.stream('myGraph', {})
YIELD nodeId, score
RETURN gds.util.asNode(nodeId).name AS name, score
ORDER BY score DESC LIMIT 10

// Multiple properties — convert once with WITH
CALL gds.pageRank.stream('myGraph', {})
YIELD nodeId, score
WITH gds.util.asNode(nodeId) AS node, score
RETURN node.name AS name, node.born AS born, score
ORDER BY score DESC LIMIT 10

Not needed for write, mutate, or stats modes — those don't return per-node data.

---

Core Algorithms

PageRank (centrality)

CALL gds.pageRank.stream('myGraph', { dampingFactor: 0.85, maxIterations: 20 })
YIELD nodeId, score
RETURN gds.util.asNode(nodeId).name AS name, score ORDER BY score DESC LIMIT 10
// score: relative influence — not absolute. Compare within same run only.
// didConverge: true means score stabilized; if false, increase maxIterations.

CALL gds.pageRank.write('myGraph', { writeProperty: 'pagerank', dampingFactor: 0.85 })
YIELD nodePropertiesWritten, ranIterations, didConverge
pr_df = gds.v2.page_rank.stream(G, damping_factor=0.85)
mutate_result = gds.v2.page_rank.mutate(G, mutate_property="pagerank", damping_factor=0.85)
write_result = gds.v2.page_rank.write(G, write_property="pagerank", damping_factor=0.85)
print(write_result.write_millis)

Louvain (community detection)

CALL gds.louvain.stream('myGraph', { relationshipWeightProperty: 'weight' })
YIELD nodeId, communityId

CALL gds.louvain.write('myGraph', { writeProperty: 'community' })
YIELD communityCount, modularity
louvain_df = gds.v2.louvain.stream(G)
write_result = gds.v2.louvain.write(G, write_property="community")
print(write_result.community_count)

Leiden is a refinement of Louvain avoiding poorly connected communities — use when community quality > raw speed. modularity in stats result: range -0.5 to 1.0. [field] Values > 0.3 often indicate meaningful community structure; > 0.7 is strong. Leiden is defined for directed and undirected graphs. Project undirected relationships when community structure is naturally symmetric.

WCC — Weakly Connected Components

Run WCC first to understand graph structure; partition disconnected graphs before expensive algorithms.

CALL gds.wcc.stream('myGraph', { minComponentSize: 10 })
YIELD nodeId, componentId

CALL gds.wcc.write('myGraph', { writeProperty: 'componentId' })
YIELD nodePropertiesWritten, componentCount
wcc_df = gds.v2.wcc.stream(G)
write_result = gds.v2.wcc.write(G, write_property="componentId")
print(write_result.node_properties_written)

Betweenness Centrality

gds.v2.betweenness_centrality.stream(G)          # identifies bottleneck/bridge nodes
gds.v2.betweenness_centrality.write(G, write_property="betweenness")

Node Similarity

Jaccard similarity from common neighbors — no node properties required.

gds.v2.node_similarity.stream(G, similarity_cutoff=0.1, top_k=10)
gds.v2.node_similarity.write(G, write_relationship_type="SIMILAR", write_property="score",
                             similarity_cutoff=0.1, top_k=10)

FastRP (node embeddings)

Fast, scalable, production ML pipelines. Set randomSeed for reproducibility.

CALL gds.fastRP.mutate('myGraph', {
  embeddingDimension: 256,
  iterationWeights: [0.0, 1.0, 1.0],
  featureProperties: ['score'],
  propertyRatio: 0.5,
  normalizationStrength: -0.5,
  randomSeed: 42,
  mutateProperty: 'embedding'
})
YIELD nodePropertiesWritten
gds.v2.fast_rp.mutate(G, embedding_dimension=256, iteration_weights=[0.0, 1.0, 1.0],
                      random_seed=42, mutate_property="embedding")
write_result = gds.v2.fast_rp.write(G, embedding_dimension=256, write_property="embedding",
                                    random_seed=42)
print(write_result.write_millis)

For ANN search over structural embeddings, after write, create a Neo4j vector index over the written property. Use neo4j-vector-index-skill.

KNN — K-Nearest Neighbors

Finds k most similar nodes per node based on node properties (typically embeddings).

CALL gds.knn.stream('myGraph', {
  nodeProperties: ['embedding'], topK: 10,
  sampleRate: 0.5, similarityCutoff: 0.7
})
YIELD node1, node2, similarity

CALL gds.knn.write('myGraph', {
  nodeProperties: ['embedding'], topK: 10,
  writeRelationshipType: 'SIMILAR', writeProperty: 'score'
})
YIELD relationshipsWritten
knn_df = gds.v2.knn.stream(G, node_properties=["embedding"], top_k=10)
gds.v2.knn.write(G, node_properties=["embedding"], top_k=10,
                 write_relationship_type="SIMILAR", write_property="score")

---

FastRP → KNN Pipeline (recommendation)

# 1. Project
G, _ = gds.v2.graph.project("myGraph", "Product",
    {"BOUGHT_TOGETHER": {"orientation": "UNDIRECTED"}})

# 2. Estimate memory
print(gds.v2.fast_rp.estimate(G, embedding_dimension=128).required_memory)

# 3. Embed
gds.v2.fast_rp.mutate(G, embedding_dimension=128, random_seed=42, mutate_property="emb")

# 4. Similarity
gds.v2.knn.write(G, node_properties=["emb"], top_k=10,
                 write_relationship_type="SIMILAR", write_property="score")

# 5. Cleanup
gds.v2.graph.drop(G)

---

Algorithm Selection

GoalAlgorithm
Influence via network linksPageRank / ArticleRank
Bottleneck / bridge nodesBetweenness Centrality
Direct connectionsDegree Centrality
Community (general, fast)Louvain
Community (higher quality)Leiden
Is graph connected?WCC (run first)
Similarity from embeddingsKNN
Similarity from neighborsNode Similarity
Shortest path (positive weights)Dijkstra / A*
k alternative pathsYen's
Fast scalable embeddingsFastRP
Feature-rich nodesGraphSAGE (gds.beta.graphSage)

Full algorithm catalog → references/algorithms.md

---

Common Errors

ErrorCauseFix
Unknown function 'gds.version'Embedded GDS plugin unavailableAGA → neo4j-aura-graph-analytics-skill; self-managed/local → install plugin
Insufficient heap memory / OOMGraph too large for available JVM heapRun gds.graph.project.estimate; increase dbms.memory.heap.max_size
Procedure not found: gds.leidenOlder or incompatible GDSCheck CALL gds.list() for available procedures; upgrade GDS or use Louvain
Node property 'X' not found after mutateProperty not projected or wrong graph nameVerify G.node_properties() includes the property; check mutate_property spelling
Graph 'myGraph' already existsLeftover projection from failed runCALL gds.graph.drop('myGraph') or gds.v2.graph.drop(G)
mutate_property already existsRe-running algorithm on same projectionDrop and re-project, or use different mutate_property name
No algorithm resultsSource/target node not in projectionVerify node labels/rel types match projection; check G.node_count()

---

Full Workflow

1. Create gds with GraphDataScience(...). 2. Verify plugin: gds.server_version() or RETURN gds.version(). 3. Estimate memory: gds.graph.project.estimate(...) and algorithm .estimate(...). 4. Project named graph with gds.v2.graph.project(...). 5. Run gds.v2.*.stream first; switch to mutate; use write only when satisfied. 6. Drop graph with gds.v2.graph.drop(G). 7. Use v1 only for endpoints missing in v2, such as plugin Cypher projection.

Built-in test datasets: gds.v2.graph.datasets.load_cora(), gds.v2.graph.datasets.load_karate_club(), gds.v2.graph.datasets.load_imdb()

---

MCP Tool Mapping

OperationMCP tool
RETURN gds.version()read-cypher
gds.pageRank.stream(...)read-cypher
gds.pageRank.write(...)write-cypher
gds.graph.drop(...)write-cypher
List available proceduresread-cypherCALL gds.list()

Before any write-cypher: show exact Cypher, expected nodes/relationships affected, and ask for confirmation. For algorithm write mode, estimate or run stats first when available.

---

References

  • references/algorithms.md — full algorithm catalog: all procedures, parameters, tiers, Cypher + Python examples
  • references/graph-projection.md — projection deep-dive: filtering, heterogeneous graphs, relationship orientation, property types
  • GDS Manual
  • Python Client Docs

---

Checklist

  • [ ] Embedded GDS plugin confirmed with gds.version() or gds.server_version()
  • [ ] Graph/algorithm memory estimated before large work
  • [ ] Python examples prefer gds.v2.*, snake_case params, typed result attributes
  • [ ] v1 APIs used only as explicit fallback
  • [ ] Projection uses native or plugin Cypher projection; no gds.graph.project.remote(...)
  • [ ] Named graph dropped after use (gds.v2.graph.drop(G) or v1 fallback)
  • [ ] Execution mode chosen: stream (inspect) → mutate (chain) → write (persist)
  • [ ] write_property/mutate_property checked for collision with existing properties
  • [ ] randomSeed set for reproducible embeddings
  • [ ] WCC run first on graphs that may be disconnected

Related skills

How it compares

Pick neo4j-gds-skill over generic Cypher skills when graph algorithm projections, memory estimation, or embedding pipelines require the GDS plugin.

FAQ

Which Neo4j GDS algorithms does neo4j-gds-skill cover?

neo4j-gds-skill covers PageRank, Louvain, WCC, Betweenness Centrality, Node Similarity, FastRP, and KNN. The skill also documents the FastRP-to-KNN recommendation pipeline and embedding write-back for vector indexes.

Does neo4j-gds-skill support memory estimation?

neo4j-gds-skill includes memory estimation guidance before large native or Cypher graph projections and algorithm runs. Developers use the graphdatascience Python client to avoid heap failures on big graphs.

What GDS execution modes does neo4j-gds-skill explain?

neo4j-gds-skill explains stream, stats, mutate, and write execution modes and when to use each. The skill helps pick the right mode for analytics previews versus persisting results to the graph.

Is Neo4j Gds Skill safe to install?

skills.sh reports 3 of 3 security scanners passed. Review the Security Audits panel on this page before installing in production.

Databasesdatabasesanalytics

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.