
Spark Version Upgrade
- 2 installs
- 134 repo stars
- Updated August 4, 2026
- openhands/extensions
Migrate Apache Spark and PySpark applications across major versions by updating build files, deprecated APIs, config properties, and SQL behavior.
About
Upgrades Apache Spark applications between major versions using a six-phase workflow covering build files, deprecated APIs, config renames, and SQL behavior changes. A developer uses it when migrating PySpark or Spark/Scala code from 2.x to 3.x or 3.x to 4.x and resolving deprecation warnings.
- Phase-by-phase workflow: inventory, build files, API, config, SQL, tests
- Covers Maven/SBT/Gradle/PySpark and 2.x->3.x and 3.x->4.x breaking changes
Spark Version Upgrade by the numbers
- 2 all-time installs (skills.sh)
- Ranked #1,759 of 2,064 Data Science & ML skills by installs in the Skillselion catalog
- Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/openhands/extensions --skill spark-version-upgradeAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 2 |
|---|---|
| repo stars | ★ 134 |
| Last updated | August 4, 2026 |
| Repository | openhands/extensions ↗ |
What it does
Migrate Apache Spark and PySpark applications across major versions by updating build files, deprecated APIs, config properties, and SQL behavior.
Files
Upgrade Apache Spark applications between major versions with a structured, phase-by-phase workflow.
When to Use
- Migrating from Spark 2.x → 3.x or Spark 3.x → 4.x
- Updating PySpark, Spark SQL, or Structured Streaming applications
- Resolving deprecation warnings before a Spark version bump
Workflow Overview
1. Inventory & Impact Analysis — Scan the codebase and assess scope 2. Build File Updates — Bump Spark/Scala/Java dependencies 3. API Migration — Replace deprecated and removed APIs 4. Configuration Migration — Update Spark config properties 5. SQL & DataFrame Migration — Fix query-level breaking changes 6. Test Validation — Compile, run tests, verify results
---
Phase 1: Inventory & Impact Analysis
Before changing any code, assess what needs to change. Read the official Apache Spark migration guide for the target version — it documents every API removal, config rename, and behavioral change per release: https://spark.apache.org/docs/latest/migration-guide.html
Checklist
- [ ] Read the migration guide section for the target Spark version
- [ ] Identify current Spark version (check
pom.xml,build.sbt,build.gradle, orrequirements.txt) - [ ] Identify target Spark version
- [ ] Search for deprecated APIs:
grep -rn 'import org.apache.spark' --include='*.scala' --include='*.java' --include='*.py' - [ ] List all Spark config properties:
grep -rn 'spark\.' --include='*.conf' --include='*.properties' --include='*.scala' --include='*.java' --include='*.py' | grep -v 'test' - [ ] Check for custom
SparkSessionorSparkContextextensions - [ ] Identify connector dependencies (Hive, Kafka, Cassandra, Delta, Iceberg)
- [ ] Document findings in
spark_upgrade_impact.md
Output
spark_upgrade_impact.md # Summary of affected files, APIs, and configs---
Phase 2: Build File Updates
Update dependency versions and resolve compilation.
Maven (pom.xml)
<!-- Update Spark version property -->
<spark.version>3.5.1</spark.version> <!-- or 4.0.0 -->
<scala.version>2.13.12</scala.version> <!-- Spark 3.x: 2.12/2.13; Spark 4.x: 2.13 -->
<!-- Update artifact IDs if Scala cross-version changed -->
<artifactId>spark-core_2.13</artifactId>
<artifactId>spark-sql_2.13</artifactId>SBT (build.sbt)
val sparkVersion = "3.5.1" // or "4.0.0"
scalaVersion := "2.13.12"
libraryDependencies += "org.apache.spark" %% "spark-core" % sparkVersion
libraryDependencies += "org.apache.spark" %% "spark-sql" % sparkVersionGradle (build.gradle)
ext {
sparkVersion = '3.5.1' // or '4.0.0'
}
dependencies {
implementation "org.apache.spark:spark-core_2.13:${sparkVersion}"
implementation "org.apache.spark:spark-sql_2.13:${sparkVersion}"
}PySpark (requirements.txt / pyproject.toml)
pyspark==3.5.1 # or 4.0.0Checklist
- [ ] Update Spark version in build file
- [ ] Update Scala version if crossing 2.12→2.13 boundary
- [ ] Update Java source/target level if required (Spark 4.x requires Java 17+)
- [ ] Update connector library versions to match new Spark version
- [ ] Resolve dependency conflicts (
mvn dependency:tree/sbt dependencyTree) - [ ] Confirm project compiles (errors at this stage are expected — they guide Phase 3)
---
Phase 3: API Migration
Replace removed and deprecated APIs. Work through compiler errors systematically.
Common Patterns
Consult the official Apache Spark migration guide for the complete list of changes for each version: https://spark.apache.org/docs/latest/migration-guide.html
SparkSession Creation (2.x → 3.x)
// BEFORE (Spark 1.x/2.x)
val sc = new SparkContext(conf)
val sqlContext = new SQLContext(sc)
// AFTER (Spark 2.x+/3.x)
val spark = SparkSession.builder()
.config(conf)
.enableHiveSupport() // if needed
.getOrCreate()
val sc = spark.sparkContextRDD to DataFrame (2.x → 3.x)
// BEFORE
rdd.toDF() // implicit from SQLContext
// AFTER
import spark.implicits._
rdd.toDF() // implicit from SparkSessionAccumulator API (2.x → 3.x)
// BEFORE
val acc = sc.accumulator(0)
// AFTER
val acc = sc.longAccumulator("name")Checklist
- [ ] Replace
SQLContext/HiveContextwithSparkSession - [ ] Replace deprecated
AccumulatorwithAccumulatorV2 - [ ] Update
DataFrame→Dataset[Row]where needed - [ ] Replace removed
RDD.mapPartitionsWithContextwithmapPartitions - [ ] Fix
SparkConfdeprecated setters - [ ] Update custom
UserDefinedFunctionregistration - [ ] Migrate
Experimental/DeveloperApiusages that were removed - [ ] Verify all compilation errors from Phase 2 are resolved
---
Phase 4: Configuration Migration
Spark renames and removes configuration properties between versions. The official migration guide documents every renamed and removed property per release: https://spark.apache.org/docs/latest/migration-guide.html
Checklist
- [ ] Rename deprecated config keys (e.g.,
spark.shuffle.file.buffer.kb→spark.shuffle.file.buffer) - [ ] Update removed configs to their replacements
- [ ] Review
spark-defaults.conf, application code, and submit scripts - [ ] Check for hardcoded config values in test fixtures
- [ ] Verify
SparkSession.builder().config(...)calls use current property names
---
Phase 5: SQL & DataFrame Migration
Spark SQL behavior changes between versions can silently alter query results.
Key Breaking Changes (2.x → 3.x)
CASTto integer no longer truncates silently — setspark.sql.ansi.enabledif neededFROMclause is required inSELECT(no moreSELECT 1)- Column resolution order changed in subqueries
spark.sql.legacy.timeParserPolicycontrols date/time parsing behavior
Key Breaking Changes (3.x → 4.x)
- ANSI mode is default (
spark.sql.ansi.enabled=true) - Stricter type coercion in comparisons
spark.sql.legacy.*flags removed
Checklist
- [ ] Audit SQL strings and DataFrame expressions for changed behavior
- [ ] Add explicit
CASTwhere implicit coercion relied on legacy behavior - [ ] Update date/time format patterns to match new parser
- [ ] Test SQL queries with representative data and compare output to pre-upgrade baseline
- [ ] Set
spark.sql.legacy.*flags temporarily if needed for phased migration
---
Phase 6: Test Validation
Checklist
- [ ] All code compiles without errors
- [ ] All existing unit tests pass
- [ ] All existing integration tests pass
- [ ] Run Spark jobs locally with sample data and compare output to pre-upgrade baseline
- [ ] No deprecation warnings remain (or are documented with a migration timeline)
- [ ] Update CI/CD pipeline to use new Spark version
- [ ] Document any
spark.sql.legacy.*flags that are set temporarily
Done When
✓ Project compiles against target Spark version ✓ All tests pass ✓ No removed APIs remain in code ✓ Configuration properties are current ✓ SQL queries produce correct results ✓ Upgrade impact documented in spark_upgrade_impact.md
.plugin.plugin{
"name": "spark-version-upgrade",
"version": "1.0.0",
"description": "Upgrade Apache Spark applications between major versions (2.x\u21923.x, 3.x\u21924.x). Covers build files, deprecated APIs, configuration changes, SQL/DataFrame updates, and test validation.",
"author": {
"name": "OpenHands",
"email": "contact@all-hands.dev"
},
"homepage": "https://github.com/OpenHands/extensions",
"repository": "https://github.com/OpenHands/extensions",
"license": "MIT",
"keywords": [
"spark",
"upgrade",
"migration",
"pyspark",
"scala",
"big-data"
]
}
spark-version-upgrade
Upgrade Apache Spark applications between major versions (2.x→3.x, 3.x→4.x). Covers build files, deprecated APIs, configuration changes, SQL/DataFrame updates, and test validation.
Triggers
This skill is activated by the following keywords:
spark upgradespark migrationspark versionupgrade sparkspark 3spark 4pyspark upgrade
Overview
This skill provides a structured, six-phase workflow for upgrading Apache Spark applications:
| Phase | Description |
|---|---|
| 1. Inventory & Impact Analysis | Scan the codebase, identify Spark usage, document scope |
| 2. Build File Updates | Bump Spark/Scala/Java versions in Maven, SBT, Gradle, or pip |
| 3. API Migration | Replace removed/deprecated APIs (SQLContext, Accumulator, etc.) |
| 4. Configuration Migration | Rename/remove deprecated Spark config properties |
| 5. SQL & DataFrame Migration | Fix breaking SQL behavior (ANSI mode, type coercion, date parsing) |
| 6. Test Validation | Compile, test, compare output to pre-upgrade baseline |
Supported Upgrade Paths
- Spark 2.x → 3.x — Major API removals (SQLContext, HiveContext, Accumulator v1), Scala 2.12/2.13
- Spark 3.x → 4.x — ANSI mode default, Java 17+ requirement, Scala 2.13 only, legacy flag removal
Languages & Build Systems
- Languages: Scala, Java, Python (PySpark)
- Build systems: Maven, SBT, Gradle, pip/uv
Reference Material
- Apache Spark Migration Guide — The official, up-to-date guide covering API removals, configuration changes, SQL behavior, PySpark, Structured Streaming, and MLlib for every Spark release
Example Usage
Ask the agent:
"Upgrade this project from Spark 2.4 to Spark 3.5"
"Migrate our PySpark codebase to Spark 4.0"
"Fix all Spark deprecation warnings in this repo"
The agent will follow the six-phase workflow, producing a spark_upgrade_impact.md document and systematically updating build files, code, configuration, and SQL queries.