
Csv Data Summarizer
- 49 installs
- 16 repo stars
- Updated November 20, 2025
- jackspace/claudeskillz
This is a copy of csv-data-summarizer by coffeefuelbump - installs and ranking accrue to the original listing.
Helps with ai & agent building tasks.
About
csv-data-summarizer is a Claude Code skill for ai & agent building. It helps solo builders move faster with AI-assisted coding.
- csv-data-summarizer
- AI & Agent Building
- AI-coding skill
Csv Data Summarizer by the numbers
- 49 all-time installs (skills.sh)
- Data as of Aug 2, 2026 (Skillselion catalog sync)
npx skills add https://github.com/jackspace/claudeskillz --skill csv-data-summarizerAdd your badge
Show developers this skill is listed on Skillselion. Paste this into your README.
| Installs | 49 |
|---|---|
| repo stars | ★ 16 |
| Last updated | November 20, 2025 |
| Repository | jackspace/claudeskillz ↗ |
What it does
Helps with ai & agent building tasks.
Files
CSV Data Summarizer
This Skill analyzes CSV files and provides comprehensive summaries with statistical insights and visualizations.
When to Use This Skill
Claude should use this Skill whenever the user:
- Uploads or references a CSV file
- Asks to summarize, analyze, or visualize tabular data
- Requests insights from CSV data
- Wants to understand data structure and quality
How It Works
⚠️ CRITICAL BEHAVIOR REQUIREMENT ⚠️
DO NOT ASK THE USER WHAT THEY WANT TO DO WITH THE DATA. DO NOT OFFER OPTIONS OR CHOICES. DO NOT SAY "What would you like me to help you with?" DO NOT LIST POSSIBLE ANALYSES.
IMMEDIATELY AND AUTOMATICALLY: 1. Run the comprehensive analysis 2. Generate ALL relevant visualizations 3. Present complete results 4. NO questions, NO options, NO waiting for user input
THE USER WANTS A FULL ANALYSIS RIGHT AWAY - JUST DO IT.
Automatic Analysis Steps:
The skill intelligently adapts to different data types and industries by inspecting the data first, then determining what analyses are most relevant.
1. Load and inspect the CSV file into pandas DataFrame 2. Identify data structure - column types, date columns, numeric columns, categories 3. Determine relevant analyses based on what's actually in the data:
- Sales/E-commerce data (order dates, revenue, products): Time-series trends, revenue analysis, product performance
- Customer data (demographics, segments, regions): Distribution analysis, segmentation, geographic patterns
- Financial data (transactions, amounts, dates): Trend analysis, statistical summaries, correlations
- Operational data (timestamps, metrics, status): Time-series, performance metrics, distributions
- Survey data (categorical responses, ratings): Frequency analysis, cross-tabulations, distributions
- Generic tabular data: Adapts based on column types found
4. Only create visualizations that make sense for the specific dataset:
- Time-series plots ONLY if date/timestamp columns exist
- Correlation heatmaps ONLY if multiple numeric columns exist
- Category distributions ONLY if categorical columns exist
- Histograms for numeric distributions when relevant
5. Generate comprehensive output automatically including:
- Data overview (rows, columns, types)
- Key statistics and metrics relevant to the data type
- Missing data analysis
- Multiple relevant visualizations (only those that apply)
- Actionable insights based on patterns found in THIS specific dataset
6. Present everything in one complete analysis - no follow-up questions
Example adaptations:
- Healthcare data with patient IDs → Focus on demographics, treatment patterns, temporal trends
- Inventory data with stock levels → Focus on quantity distributions, reorder patterns, SKU analysis
- Web analytics with timestamps → Focus on traffic patterns, conversion metrics, time-of-day analysis
- Survey responses → Focus on response distributions, demographic breakdowns, sentiment patterns
Behavior Guidelines
✅ CORRECT APPROACH - SAY THIS:
- "I'll analyze this data comprehensively right now."
- "Here's the complete analysis with visualizations:"
- "I've identified this as [type] data and generated relevant insights:"
- Then IMMEDIATELY show the full analysis
✅ DO:
- Immediately run the analysis script
- Generate ALL relevant charts automatically
- Provide complete insights without being asked
- Be thorough and complete in first response
- Act decisively without asking permission
❌ NEVER SAY THESE PHRASES:
- "What would you like to do with this data?"
- "What would you like me to help you with?"
- "Here are some common options:"
- "Let me know what you'd like help with"
- "I can create a comprehensive analysis if you'd like!"
- Any sentence ending with "?" asking for user direction
- Any list of options or choices
- Any conditional "I can do X if you want"
❌ FORBIDDEN BEHAVIORS:
- Asking what the user wants
- Listing options for the user to choose from
- Waiting for user direction before analyzing
- Providing partial analysis that requires follow-up
- Describing what you COULD do instead of DOING it
Usage
The Skill provides a Python function summarize_csv(file_path) that:
- Accepts a path to a CSV file
- Returns a comprehensive text summary with statistics
- Generates multiple visualizations automatically based on data structure
Example Prompts
"Here's sales_data.csv. Can you summarize this file?""Analyze this customer data CSV and show me trends."
"What insights can you find in orders.csv?"Example Output
Dataset Overview
- 5,000 rows × 8 columns
- 3 numeric columns, 1 date column
Summary Statistics
- Average order value: $58.2
- Standard deviation: $12.4
- Missing values: 2% (100 cells)
Insights
- Sales show upward trend over time
- Peak activity in Q4
(Attached: trend plot)
Files
analyze.py- Core analysis logicrequirements.txt- Python dependenciesresources/sample.csv- Example dataset for testingresources/README.md- Additional documentation
Notes
- Automatically detects date columns (columns containing 'date' in name)
- Handles missing data gracefully
- Generates visualizations only when date columns are present
- All numeric columns are included in statistical summary
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
from pathlib import Path
def summarize_csv(file_path):
"""
Comprehensively analyzes a CSV file and generates multiple visualizations.
Args:
file_path (str): Path to the CSV file
Returns:
str: Formatted comprehensive analysis of the dataset
"""
df = pd.read_csv(file_path)
summary = []
charts_created = []
# Basic info
summary.append("=" * 60)
summary.append("📊 DATA OVERVIEW")
summary.append("=" * 60)
summary.append(f"Rows: {df.shape[0]:,} | Columns: {df.shape[1]}")
summary.append(f"\nColumns: {', '.join(df.columns.tolist())}")
# Data types
summary.append(f"\n📋 DATA TYPES:")
for col, dtype in df.dtypes.items():
summary.append(f" • {col}: {dtype}")
# Missing data analysis
missing = df.isnull().sum().sum()
missing_pct = (missing / (df.shape[0] * df.shape[1])) * 100
summary.append(f"\n🔍 DATA QUALITY:")
if missing:
summary.append(f"Missing values: {missing:,} ({missing_pct:.2f}% of total data)")
summary.append("Missing by column:")
for col in df.columns:
col_missing = df[col].isnull().sum()
if col_missing > 0:
col_pct = (col_missing / len(df)) * 100
summary.append(f" • {col}: {col_missing:,} ({col_pct:.1f}%)")
else:
summary.append("✓ No missing values - dataset is complete!")
# Numeric analysis
numeric_cols = df.select_dtypes(include='number').columns.tolist()
if numeric_cols:
summary.append(f"\n📈 NUMERICAL ANALYSIS:")
summary.append(str(df[numeric_cols].describe()))
# Correlations if multiple numeric columns
if len(numeric_cols) > 1:
summary.append(f"\n🔗 CORRELATIONS:")
corr_matrix = df[numeric_cols].corr()
summary.append(str(corr_matrix))
# Create correlation heatmap
plt.figure(figsize=(10, 8))
sns.heatmap(corr_matrix, annot=True, cmap='coolwarm', center=0,
square=True, linewidths=1)
plt.title('Correlation Heatmap')
plt.tight_layout()
plt.savefig('correlation_heatmap.png', dpi=150)
plt.close()
charts_created.append('correlation_heatmap.png')
# Categorical analysis
categorical_cols = df.select_dtypes(include=['object']).columns.tolist()
categorical_cols = [c for c in categorical_cols if 'id' not in c.lower()]
if categorical_cols:
summary.append(f"\n📊 CATEGORICAL ANALYSIS:")
for col in categorical_cols[:5]: # Limit to first 5
value_counts = df[col].value_counts()
summary.append(f"\n{col}:")
for val, count in value_counts.head(10).items():
pct = (count / len(df)) * 100
summary.append(f" • {val}: {count:,} ({pct:.1f}%)")
# Time series analysis
date_cols = [c for c in df.columns if 'date' in c.lower() or 'time' in c.lower()]
if date_cols:
summary.append(f"\n📅 TIME SERIES ANALYSIS:")
date_col = date_cols[0]
df[date_col] = pd.to_datetime(df[date_col], errors='coerce')
date_range = df[date_col].max() - df[date_col].min()
summary.append(f"Date range: {df[date_col].min()} to {df[date_col].max()}")
summary.append(f"Span: {date_range.days} days")
# Create time-series plots for numeric columns
if numeric_cols:
fig, axes = plt.subplots(min(3, len(numeric_cols)), 1,
figsize=(12, 4 * min(3, len(numeric_cols))))
if len(numeric_cols) == 1:
axes = [axes]
for idx, num_col in enumerate(numeric_cols[:3]):
ax = axes[idx] if len(numeric_cols) > 1 else axes[0]
daily_data = df.groupby(date_col)[num_col].agg(['mean', 'sum', 'count'])
daily_data['mean'].plot(ax=ax, label='Average', linewidth=2)
ax.set_title(f'{num_col} Over Time')
ax.set_xlabel('Date')
ax.set_ylabel(num_col)
ax.legend()
ax.grid(True, alpha=0.3)
plt.tight_layout()
plt.savefig('time_series_analysis.png', dpi=150)
plt.close()
charts_created.append('time_series_analysis.png')
# Distribution plots for numeric columns
if numeric_cols:
n_cols = min(4, len(numeric_cols))
fig, axes = plt.subplots(2, 2, figsize=(12, 10))
axes = axes.flatten()
for idx, col in enumerate(numeric_cols[:4]):
axes[idx].hist(df[col].dropna(), bins=30, edgecolor='black', alpha=0.7)
axes[idx].set_title(f'Distribution of {col}')
axes[idx].set_xlabel(col)
axes[idx].set_ylabel('Frequency')
axes[idx].grid(True, alpha=0.3)
# Hide unused subplots
for idx in range(len(numeric_cols[:4]), 4):
axes[idx].set_visible(False)
plt.tight_layout()
plt.savefig('distributions.png', dpi=150)
plt.close()
charts_created.append('distributions.png')
# Categorical distributions
if categorical_cols:
fig, axes = plt.subplots(2, 2, figsize=(14, 10))
axes = axes.flatten()
for idx, col in enumerate(categorical_cols[:4]):
value_counts = df[col].value_counts().head(10)
axes[idx].barh(range(len(value_counts)), value_counts.values)
axes[idx].set_yticks(range(len(value_counts)))
axes[idx].set_yticklabels(value_counts.index)
axes[idx].set_title(f'Top Values in {col}')
axes[idx].set_xlabel('Count')
axes[idx].grid(True, alpha=0.3, axis='x')
# Hide unused subplots
for idx in range(len(categorical_cols[:4]), 4):
axes[idx].set_visible(False)
plt.tight_layout()
plt.savefig('categorical_distributions.png', dpi=150)
plt.close()
charts_created.append('categorical_distributions.png')
# Summary of visualizations
if charts_created:
summary.append(f"\n📊 VISUALIZATIONS CREATED:")
for chart in charts_created:
summary.append(f" ✓ {chart}")
summary.append("\n" + "=" * 60)
summary.append("✅ COMPREHENSIVE ANALYSIS COMPLETE")
summary.append("=" * 60)
return "\n".join(summary)
if __name__ == "__main__":
# Test with sample data
import sys
if len(sys.argv) > 1:
file_path = sys.argv[1]
else:
file_path = "resources/sample.csv"
print(summarize_csv(file_path))
<div align="center">
 
 
</div>
---
📊 CSV Data Summarizer - Claude Skill
A powerful Claude Skill that automatically analyzes CSV files and generates comprehensive insights with visualizations. Upload any CSV and get instant, intelligent analysis without being asked what you want!
<div align="center">
  
</div>
🚀 Features
- 🤖 Intelligent & Adaptive - Automatically detects data type (sales, customer, financial, survey, etc.) and applies relevant analysis
- 📈 Comprehensive Analysis - Generates statistics, correlations, distributions, and trends
- 🎨 Auto Visualizations - Creates multiple charts based on what's in your data:
- Time-series plots for date-based data
- Correlation heatmaps for numeric relationships
- Distribution histograms
- Categorical breakdowns
- ⚡ Proactive - No questions asked! Just upload CSV and get complete analysis immediately
- 🔍 Data Quality Checks - Automatically detects and reports missing values
- 📊 Multi-Industry Support - Adapts to e-commerce, healthcare, finance, operations, surveys, and more
📥 Quick Download
<div align="center">
Get Started in 2 Steps
1️⃣ Download the Skill 
2️⃣ Try the Demo Data 
</div>
---
📦 What's Included
csv-data-summarizer-claude-skill/
├── SKILL.md # Claude Skill definition
├── analyze.py # Comprehensive analysis engine
├── requirements.txt # Python dependencies
├── examples/
│ └── showcase_financial_pl_data.csv # Demo P&L financial dataset (15 months, 25 metrics)
└── resources/
├── sample.csv # Example dataset
└── README.md # Usage documentation🎯 How It Works
1. Upload any CSV file to Claude.ai 2. Skill activates automatically when CSV is detected 3. Analysis runs immediately - inspects data structure and adapts 4. Results delivered - Complete analysis with multiple visualizations
No prompting needed. No options to choose. Just instant, comprehensive insights!
📥 Installation
For Claude.ai Users
1. Download the latest release: `csv-data-summarizer.zip` 2. Go to Claude.ai → Settings → Capabilities → Skills 3. Upload the zip file 4. Enable the skill 5. Done! Upload any CSV and watch it work ✨
For Developers
git clone git@github.com:coffeefuelbump/csv-data-summarizer-claude-skill.git
cd csv-data-summarizer-claude-skill
pip install -r requirements.txt📊 Sample Dataset Highlights
The included demo CSV contains 15 months of P&L data with:
- 3 product lines (SaaS, Enterprise, Services)
- 25 financial metrics including revenue, expenses, margins, CAC, LTV
- Quarterly trends showing business growth
- Perfect for showcasing time-series analysis, correlations, and financial insights
🎨 Example Use Cases
- 📊 Sales Data → Revenue trends, product performance, regional analysis
- 👥 Customer Data → Demographics, segmentation, geographic patterns
- 💰 Financial Data → Transaction analysis, trend detection, correlations
- ⚙️ Operational Data → Performance metrics, time-series analysis
- 📋 Survey Data → Response distributions, cross-tabulations
🛠️ Technical Details
Dependencies:
- Python 3.8+
- pandas 2.0+
- matplotlib 3.7+
- seaborn 0.12+
Visualizations Generated:
- Time-series trend plots
- Correlation heatmaps
- Distribution histograms
- Categorical bar charts
📝 Example Output
============================================================
📊 DATA OVERVIEW
============================================================
Rows: 100 | Columns: 15
📋 DATA TYPES:
• order_date: object
• total_revenue: float64
• customer_segment: object
...
🔍 DATA QUALITY:
✓ No missing values - dataset is complete!
📈 NUMERICAL ANALYSIS:
[Summary statistics for all numeric columns]
🔗 CORRELATIONS:
[Correlation matrix showing relationships]
📅 TIME SERIES ANALYSIS:
Date range: 2024-01-05 to 2024-04-11
Span: 97 days
📊 VISUALIZATIONS CREATED:
✓ correlation_heatmap.png
✓ time_series_analysis.png
✓ distributions.png
✓ categorical_distributions.png🌟 Connect & Learn More
<div align="center">




</div>
🤝 Contributing
Contributions are welcome! Feel free to:
- Report bugs
- Suggest new features
- Submit pull requests
- Share your use cases
📄 License
MIT License - feel free to use this skill for personal or commercial projects!
🙏 Acknowledgments
Built for the Claude Skills platform by Anthropic.
---
<div align="center">
Made with ❤️ for the AI community
⭐ Star this repo if you find it useful!
</div>
pandas>=2.0.0
matplotlib>=3.7.0
seaborn>=0.12.0
{
"sections": {
"Notes": "- Automatically detects date columns (columns containing 'date' in name)\r\n- Handles missing data gracefully\r\n- Generates visualizations only when date columns are present\r\n- All numeric columns are included in statistical summary",
"Files": "- `analyze.py` - Core analysis logic\r\n- `requirements.txt` - Python dependencies\r\n- `resources/sample.csv` - Example dataset for testing\r\n- `resources/README.md` - Additional documentation",
"⚠️ CRITICAL BEHAVIOR REQUIREMENT ⚠️": "**DO NOT ASK THE USER WHAT THEY WANT TO DO WITH THE DATA.**\r\n**DO NOT OFFER OPTIONS OR CHOICES.**\r\n**DO NOT SAY \"What would you like me to help you with?\"**\r\n**DO NOT LIST POSSIBLE ANALYSES.**\r\n\r\n**IMMEDIATELY AND AUTOMATICALLY:**\r\n1. Run the comprehensive analysis\r\n2. Generate ALL relevant visualizations\r\n3. Present complete results\r\n4. NO questions, NO options, NO waiting for user input\r\n\r\n**THE USER WANTS A FULL ANALYSIS RIGHT AWAY - JUST DO IT.**\r\n\r\n### Automatic Analysis Steps:\r\n\r\n**The skill intelligently adapts to different data types and industries by inspecting the data first, then determining what analyses are most relevant.**\r\n\r\n1. **Load and inspect** the CSV file into pandas DataFrame\r\n2. **Identify data structure** - column types, date columns, numeric columns, categories\r\n3. **Determine relevant analyses** based on what's actually in the data:\r\n - **Sales/E-commerce data** (order dates, revenue, products): Time-series trends, revenue analysis, product performance\r\n - **Customer data** (demographics, segments, regions): Distribution analysis, segmentation, geographic patterns\r\n - **Financial data** (transactions, amounts, dates): Trend analysis, statistical summaries, correlations\r\n - **Operational data** (timestamps, metrics, status): Time-series, performance metrics, distributions\r\n - **Survey data** (categorical responses, ratings): Frequency analysis, cross-tabulations, distributions\r\n - **Generic tabular data**: Adapts based on column types found\r\n\r\n4. **Only create visualizations that make sense** for the specific dataset:\r\n - Time-series plots ONLY if date/timestamp columns exist\r\n - Correlation heatmaps ONLY if multiple numeric columns exist\r\n - Category distributions ONLY if categorical columns exist\r\n - Histograms for numeric distributions when relevant\r\n \r\n5. **Generate comprehensive output** automatically including:\r\n - Data overview (rows, columns, types)\r\n - Key statistics and metrics relevant to the data type\r\n - Missing data analysis\r\n - Multiple relevant visualizations (only those that apply)\r\n - Actionable insights based on patterns found in THIS specific dataset\r\n \r\n6. **Present everything** in one complete analysis - no follow-up questions\r\n\r\n**Example adaptations:**\r\n- Healthcare data with patient IDs → Focus on demographics, treatment patterns, temporal trends\r\n- Inventory data with stock levels → Focus on quantity distributions, reorder patterns, SKU analysis \r\n- Web analytics with timestamps → Focus on traffic patterns, conversion metrics, time-of-day analysis\r\n- Survey responses → Focus on response distributions, demographic breakdowns, sentiment patterns\r\n\r\n### Behavior Guidelines\r\n\r\n✅ **CORRECT APPROACH - SAY THIS:**\r\n- \"I'll analyze this data comprehensively right now.\"\r\n- \"Here's the complete analysis with visualizations:\"\r\n- \"I've identified this as [type] data and generated relevant insights:\"\r\n- Then IMMEDIATELY show the full analysis\r\n\r\n✅ **DO:**\r\n- Immediately run the analysis script\r\n- Generate ALL relevant charts automatically\r\n- Provide complete insights without being asked\r\n- Be thorough and complete in first response\r\n- Act decisively without asking permission\r\n\r\n❌ **NEVER SAY THESE PHRASES:**\r\n- \"What would you like to do with this data?\"\r\n- \"What would you like me to help you with?\"\r\n- \"Here are some common options:\"\r\n- \"Let me know what you'd like help with\"\r\n- \"I can create a comprehensive analysis if you'd like!\"\r\n- Any sentence ending with \"?\" asking for user direction\r\n- Any list of options or choices\r\n- Any conditional \"I can do X if you want\"\r\n\r\n❌ **FORBIDDEN BEHAVIORS:**\r\n- Asking what the user wants\r\n- Listing options for the user to choose from\r\n- Waiting for user direction before analyzing\r\n- Providing partial analysis that requires follow-up\r\n- Describing what you COULD do instead of DOING it\r\n\r\n### Usage\r\n\r\nThe Skill provides a Python function `summarize_csv(file_path)` that:\r\n- Accepts a path to a CSV file\r\n- Returns a comprehensive text summary with statistics\r\n- Generates multiple visualizations automatically based on data structure\r\n\r\n### Example Prompts\r\n\r\n> \"Here's `sales_data.csv`. Can you summarize this file?\"\r\n\r\n> \"Analyze this customer data CSV and show me trends.\"\r\n\r\n> \"What insights can you find in `orders.csv`?\"\r\n\r\n### Example Output\r\n\r\n**Dataset Overview**\r\n- 5,000 rows × 8 columns \r\n- 3 numeric columns, 1 date column \r\n\r\n**Summary Statistics**\r\n- Average order value: $58.2 \r\n- Standard deviation: $12.4\r\n- Missing values: 2% (100 cells)\r\n\r\n**Insights**\r\n- Sales show upward trend over time\r\n- Peak activity in Q4\r\n*(Attached: trend plot)*",
"When to Use This Skill": "Claude should use this Skill whenever the user:\r\n- Uploads or references a CSV file\r\n- Asks to summarize, analyze, or visualize tabular data\r\n- Requests insights from CSV data\r\n- Wants to understand data structure and quality",
"How It Works": ""
},
"content": "This Skill analyzes CSV files and provides comprehensive summaries with statistical insights and visualizations.",
"id": "csv-data-summarizer-claude-skill_coffeefuelbump",
"name": "csv-data-summarizer",
"description": "Analyzes CSV files, generates summary stats, and plots quick visualizations using Python and pandas."
}---
name: csv-data-summarizer
description: Analyzes CSV files, generates summary stats, and plots quick visualizations using Python and pandas.
metadata:
version: 2.1.0
dependencies: python>=3.8, pandas>=2.0.0, matplotlib>=3.7.0, seaborn>=0.12.0
---
# CSV Data Summarizer
This Skill analyzes CSV files and provides comprehensive summaries with statistical insights and visualizations.
## When to Use This Skill
Claude should use this Skill whenever the user:
- Uploads or references a CSV file
- Asks to summarize, analyze, or visualize tabular data
- Requests insights from CSV data
- Wants to understand data structure and quality
## How It Works
## ⚠️ CRITICAL BEHAVIOR REQUIREMENT ⚠️
**DO NOT ASK THE USER WHAT THEY WANT TO DO WITH THE DATA.**
**DO NOT OFFER OPTIONS OR CHOICES.**
**DO NOT SAY "What would you like me to help you with?"**
**DO NOT LIST POSSIBLE ANALYSES.**
**IMMEDIATELY AND AUTOMATICALLY:**
1. Run the comprehensive analysis
2. Generate ALL relevant visualizations
3. Present complete results
4. NO questions, NO options, NO waiting for user input
**THE USER WANTS A FULL ANALYSIS RIGHT AWAY - JUST DO IT.**
### Automatic Analysis Steps:
**The skill intelligently adapts to different data types and industries by inspecting the data first, then determining what analyses are most relevant.**
1. **Load and inspect** the CSV file into pandas DataFrame
2. **Identify data structure** - column types, date columns, numeric columns, categories
3. **Determine relevant analyses** based on what's actually in the data:
- **Sales/E-commerce data** (order dates, revenue, products): Time-series trends, revenue analysis, product performance
- **Customer data** (demographics, segments, regions): Distribution analysis, segmentation, geographic patterns
- **Financial data** (transactions, amounts, dates): Trend analysis, statistical summaries, correlations
- **Operational data** (timestamps, metrics, status): Time-series, performance metrics, distributions
- **Survey data** (categorical responses, ratings): Frequency analysis, cross-tabulations, distributions
- **Generic tabular data**: Adapts based on column types found
4. **Only create visualizations that make sense** for the specific dataset:
- Time-series plots ONLY if date/timestamp columns exist
- Correlation heatmaps ONLY if multiple numeric columns exist
- Category distributions ONLY if categorical columns exist
- Histograms for numeric distributions when relevant
5. **Generate comprehensive output** automatically including:
- Data overview (rows, columns, types)
- Key statistics and metrics relevant to the data type
- Missing data analysis
- Multiple relevant visualizations (only those that apply)
- Actionable insights based on patterns found in THIS specific dataset
6. **Present everything** in one complete analysis - no follow-up questions
**Example adaptations:**
- Healthcare data with patient IDs → Focus on demographics, treatment patterns, temporal trends
- Inventory data with stock levels → Focus on quantity distributions, reorder patterns, SKU analysis
- Web analytics with timestamps → Focus on traffic patterns, conversion metrics, time-of-day analysis
- Survey responses → Focus on response distributions, demographic breakdowns, sentiment patterns
### Behavior Guidelines
✅ **CORRECT APPROACH - SAY THIS:**
- "I'll analyze this data comprehensively right now."
- "Here's the complete analysis with visualizations:"
- "I've identified this as [type] data and generated relevant insights:"
- Then IMMEDIATELY show the full analysis
✅ **DO:**
- Immediately run the analysis script
- Generate ALL relevant charts automatically
- Provide complete insights without being asked
- Be thorough and complete in first response
- Act decisively without asking permission
❌ **NEVER SAY THESE PHRASES:**
- "What would you like to do with this data?"
- "What would you like me to help you with?"
- "Here are some common options:"
- "Let me know what you'd like help with"
- "I can create a comprehensive analysis if you'd like!"
- Any sentence ending with "?" asking for user direction
- Any list of options or choices
- Any conditional "I can do X if you want"
❌ **FORBIDDEN BEHAVIORS:**
- Asking what the user wants
- Listing options for the user to choose from
- Waiting for user direction before analyzing
- Providing partial analysis that requires follow-up
- Describing what you COULD do instead of DOING it
### Usage
The Skill provides a Python function `summarize_csv(file_path)` that:
- Accepts a path to a CSV file
- Returns a comprehensive text summary with statistics
- Generates multiple visualizations automatically based on data structure
### Example Prompts
> "Here's `sales_data.csv`. Can you summarize this file?"
> "Analyze this customer data CSV and show me trends."
> "What insights can you find in `orders.csv`?"
### Example Output
**Dataset Overview**
- 5,000 rows × 8 columns
- 3 numeric columns, 1 date column
**Summary Statistics**
- Average order value: $58.2
- Standard deviation: $12.4
- Missing values: 2% (100 cells)
**Insights**
- Sales show upward trend over time
- Peak activity in Q4
*(Attached: trend plot)*
## Files
- `analyze.py` - Core analysis logic
- `requirements.txt` - Python dependencies
- `resources/sample.csv` - Example dataset for testing
- `resources/README.md` - Additional documentation
## Notes
- Automatically detects date columns (columns containing 'date' in name)
- Handles missing data gracefully
- Generates visualizations only when date columns are present
- All numeric columns are included in statistical summary