Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
secondsky avatar

Sap Hana Ml

  • 330 installs
  • 399 repo stars
  • Updated August 4, 2026
  • secondsky/sap-skills

Build in-database ML with SAP HANA: PAL/APL algorithms, model training in SQLScript, scoring pipelines, and embedding predictions beside transactional data.

About

Covers SAP HANA machine learning with PAL/APL, SQLScript pipelines, in-database training, and embedded scoring next to transactional ERP data. Suited to SaaS and API solutions needing governed, low-latency predictions without exporting sensitive data to external ML platforms.

  • SAP HANA ML (PAL/APL)
  • In-database model training
  • SQLScript scoring pipelines
  • Embedded ERP predictions
  • Low-latency analytics on HANA

Sap Hana Ml by the numbers

  • 330 all-time installs (skills.sh)
  • +29 installs in the week ending Aug 5, 2026 (Skillselion tracking)
  • Ranked #563 of 2,064 Data Science & ML skills by installs in the Skillselion catalog
  • Data as of Aug 5, 2026 (Skillselion catalog sync)
npx skills add https://github.com/secondsky/sap-skills --skill sap-hana-ml

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs330
repo stars399
Last updatedAugust 4, 2026
Repositorysecondsky/sap-skills

What it does

Build in-database ML with SAP HANA: PAL/APL algorithms, model training in SQLScript, scoring pipelines, and embedding predictions beside transactional data.

Files

SKILL.mdMarkdownGitHub ↗

SAP HANA ML Python Client (hana-ml)

Related Skills

  • sap-dependency-security: Use for secure dependency pinning and upgrade workflows in Python/auxiliary tooling used alongside HANA ML stacks

When to Use This Skill

Use this skill when building machine learning workflows with the hana-ml Python client, using PAL/APL algorithms, querying HANA DataFrames, training or scoring models in-database, using AutoML, visualizing model output, or troubleshooting Python-to-HANA ML connections.

Common Issues

IssueFirst check
Connection failsVerify HANA host, port, TLS/encryption, user privileges, and network allowlists.
PAL/APL algorithm missingConfirm the HANA system has the required AFL/PAL/APL libraries installed and licensed.
DataFrame collection is slowPush filtering/projection into HANA and avoid collecting large frames into Python.

Package Version: 2.22.241011 Last Verified: 2025-11-27

Table of Contents

---

Installation & Setup

pip install hana-ml

Requirements: Python 3.8+, SAP HANA 2.0 SPS03+ or SAP HANA Cloud

---

Quick Start

Connection & DataFrame

from hana_ml import ConnectionContext

# Connect
conn = ConnectionContext(
    address='<hostname>',
    port=443,
    user='<username>',
    password='<password>',
    encrypt=True
)

# Create DataFrame
df = conn.table('MY_TABLE', schema='MY_SCHEMA')
print(f"Shape: {df.shape}")
df.head(10).collect()

PAL Classification

from hana_ml.algorithms.pal.unified_classification import UnifiedClassification

# Train model
clf = UnifiedClassification(func='RandomDecisionTree')
clf.fit(train_df, features=['F1', 'F2', 'F3'], label='TARGET')

# Predict & evaluate
predictions = clf.predict(test_df, features=['F1', 'F2', 'F3'])
score = clf.score(test_df, features=['F1', 'F2', 'F3'], label='TARGET')

APL AutoML

from hana_ml.algorithms.apl.classification import AutoClassifier

# Automated classification
auto_clf = AutoClassifier()
auto_clf.fit(train_df, label='TARGET')
predictions = auto_clf.predict(test_df)

Model Persistence

from hana_ml.model_storage import ModelStorage

ms = ModelStorage(conn)
clf.name = 'MY_CLASSIFIER'
ms.save_model(model=clf, if_exists='replace')

---

Core Libraries

PAL (Predictive Analysis Library)

  • 100+ algorithms executed in-database
  • Categories: Classification, Regression, Clustering, Time Series, Preprocessing
  • Key classes: UnifiedClassification, UnifiedRegression, KMeans, ARIMA
  • See: references/PAL_ALGORITHMS.md for complete list

APL (Automated Predictive Library)

  • AutoML capabilities with automatic feature engineering
  • Key classes: AutoClassifier, AutoRegressor, GradientBoostingClassifier
  • See: references/APL_ALGORITHMS.md for details

DataFrames

  • Lazy evaluation - builds SQL until collect() called
  • In-database processing for optimal performance
  • See: references/DATAFRAME_REFERENCE.md for complete API

Visualizers

  • EDA plots, model explanations, metrics
  • SHAP integration for model interpretability
  • See: references/VISUALIZERS.md for 14 visualization modules

---

Common Patterns

Train-Test Split

from hana_ml.algorithms.pal.partition import train_test_val_split

train, test, val = train_test_val_split(
    data=df,
    training_percentage=0.7,
    testing_percentage=0.2,
    validation_percentage=0.1
)

Feature Importance

# APL models
importance = auto_clf.get_feature_importances()

# PAL models
from hana_ml.algorithms.pal.preprocessing import FeatureSelection
fs = FeatureSelection()
fs.fit(train_df, features=features, label='TARGET')

Pipeline

from hana_ml.algorithms.pal.pipeline import Pipeline
from hana_ml.algorithms.pal.preprocessing import Imputer, FeatureNormalizer

pipeline = Pipeline([
    ('imputer', Imputer(strategy='mean')),
    ('normalizer', FeatureNormalizer()),
    ('classifier', UnifiedClassification(func='RandomDecisionTree'))
])

---

Best Practices

1. Use lazy evaluation - Operations build SQL without execution until collect() 2. Leverage in-database processing - Keep data in HANA for performance 3. Use Unified interfaces - Consistent APIs across algorithms 4. Save models - Use ModelStorage for persistence 5. Explain predictions - Use SHAP explainers for interpretability 6. Monitor AutoML - Use PipelineProgressStatusMonitor for long-running jobs

---

Bundled Resources

Reference Files

  • `references/DATAFRAME_REFERENCE.md` (479 lines)
  • ConnectionContext API, DataFrame operations, SQL generation
  • `references/PAL_ALGORITHMS.md` (869 lines)
  • Complete PAL algorithm reference (100+ algorithms)
  • Classification, Regression, Clustering, Time Series, Preprocessing
  • `references/APL_ALGORITHMS.md` (534 lines)
  • AutoML capabilities, automated feature engineering
  • AutoClassifier, AutoRegressor, GradientBoosting classes
  • `references/VISUALIZERS.md` (704 lines)
  • 14 visualization modules (EDA, SHAP, metrics, time series)
  • Plot types, configuration, export options
  • `references/SUPPORTING_MODULES.md` (626 lines)
  • Model storage, spatial analytics, graph algorithms
  • Text mining, statistics, error handling

---

Error Handling

from hana_ml.ml_exceptions import Error

try:
    clf.fit(train_df, features=features, label='TARGET')
except Error as e:
    print(f"HANA ML Error: {e}")

---

Documentation

Related skills

Data Science & MLdatabasespipelines

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.