Overview

The preprocess summary command is your data exploration powerhouse. It generates comprehensive summary statistics that help you understand your dataset’s structure, quality, and characteristics before and after preprocessing.

For Data Scientists and Analysts, this command provides the insights you need to:

  • Assess data quality
  • Identify missing values
  • Understand distributions
  • Detect outliers
  • Validate preprocessing results

Basic Usage

Quick Start

Generate a summary for your dataset:

  preprocess summary --data ./my_dataset.csv
  

This creates a Summaryfile.toml file containing detailed statistics for each column.

Command Reference

  preprocess summary [flags]
  
FlagShorthandDefaultDescription
--data-dPath to the dataset file
--prepfile-fPrepfile.tomlUse a Prepfile for data specifications
--sep-s,CSV separator
--dsep-m.Decimal separator
--encoding-eutf-8File encoding
--output-oSummaryfile.tomlOutput filename
--html-tGenerate HTML report instead of TOML
--no-browser-bDon’t automatically open browser for HTML
--excludeExclude specific columns from summary

Output Formats

TOML Format (Default)

The default output is a TOML file containing structured summary statistics:

  preprocess summary --data ./data.csv --output stats.toml
  

This format is ideal for:

  • Programmatic access to statistics
  • Version control (can track changes in data characteristics)
  • Integration with other tools
  • Automated reporting

HTML Format

Generate a beautiful, interactive HTML report:

  preprocess summary --data ./data.csv --html --output report.html
  

The HTML report provides:

  • Color-coded visualizations
  • Sortable tables
  • Interactive exploration
  • Professional presentation
  • Easy sharing with non-technical stakeholders

What Statistics Are Generated?

For All Columns

StatisticDescriptionUse Case
rows_countNumber of non-null valuesIdentify missing data extent
typeColumn data typeUnderstand data structure
missingCount of missing valuesAssess data completeness

For Numeric Columns

StatisticDescriptionUse Case
minMinimum valueIdentify range, detect outliers
maxMaximum valueIdentify range, detect outliers
meanArithmetic meanUnderstand central tendency
medianMedian valueUnderstand central tendency (robust to outliers)
stdStandard deviationUnderstand variability
q1First quartile (25th percentile)Understand distribution
q3Third quartile (75th percentile)Understand distribution
modeMost frequent valueIdentify common values

For Text Columns

StatisticDescriptionUse Case
modality_countNumber of unique valuesAssess cardinality
top_valueMost frequent valueIdentify common categories
top_frequencyFrequency of most common valueUnderstand distribution
min_lengthMinimum string lengthIdentify formatting issues
max_lengthMaximum string lengthIdentify formatting issues
avg_lengthAverage string lengthUnderstand data characteristics

Practical Examples

Example 1: Basic Summary Generation

  preprocess summary --data ./sales_data.csv
  

This creates Summaryfile.toml with statistics for all columns.

Example 2: Generate HTML Report

  preprocess summary --data ./customer_data.csv --html
  

This creates customer_data_report.html and opens it in your default browser.

Example 3: Save to Specific File

  preprocess summary --data ./data.csv --output analysis_stats.toml
  

Example 4: Exclude Columns

Exclude sensitive or irrelevant columns from the summary:

  preprocess summary --data ./employees.csv \
  --exclude salary \
  --exclude ssn \
  --exclude password
  

Example 5: Use Prepfile Specifications

Use the data specifications from your Prepfile:

  preprocess summary --prepfile ./config/my_prep.toml
  

This is useful when you want to use the same data reading configuration as your preprocessing pipeline.

Example 6: European Format Data

  preprocess summary --data ./european_data.csv --sep ";" --dsep ","
  

Real-World Scenarios

Scenario 1: Initial Data Exploration

You’ve just received a new dataset and want to understand it:

  # First, skim the data
preprocess skim --data ./new_dataset.csv

# Then generate detailed statistics
preprocess summary --data ./new_dataset.csv --html --output exploration.html
  

Open the HTML report to:

  • See the complete structure of your data
  • Identify columns with missing values
  • Understand distributions
  • Spot potential data quality issues

Scenario 2: Pre- and Post-Processing Comparison

Compare your data before and after preprocessing:

  # Generate summary of original data
preprocess summary --data ./original_data.csv --output before_summary.toml

# Run preprocessing
preprocess run --file my_prep.toml

# Generate summary of cleaned data
preprocess summary --data ./cleaned_data.csv --output after_summary.toml
  

Now you can compare the two TOML files to see how your preprocessing affected the data.

Scenario 3: Automated Data Quality Reporting

Create a script to generate regular data quality reports:

  #!/bin/bash
# generate_quality_report.sh

DATE=$(date +%Y%m%d)

preprocess summary \
  --data ./daily_data.csv \
  --html \
  --output ./reports/quality_report_$DATE.html \
  --no-browser

echo "Data quality report generated: ./reports/quality_report_$DATE.html"
  

Scenario 4: Batch Processing Multiple Files

Generate summaries for multiple datasets:

  #!/bin/bash
# batch_summary.sh

DATASETS=("customers.csv" "products.csv" "orders.csv")

for dataset in "${DATASETS[@]}"; do
  echo "Processing $dataset..."
  preprocess summary --data ./$dataset --output summaries/${dataset%.csv}_summary.toml
  echo "Summary created for $dataset"
done
  

Understanding the Output

TOML Output Structure

The TOML output has this structure:

  [data]
filename = './my_dataset.csv'
csv_separator = ','
decimal_separator = '.'
encoding = 'utf-8'

[data_summary]
rows_count = 10000
columns_count = 15
numeric_columns = 10
string_columns = 5

[[columns]]
name = 'age'
type = 'numeric'
rows_count = 9500
missing = 500
min = 18.0
max = 85.0
mean = 42.5
median = 41.0
std = 12.3
q1 = 30.0
q3 = 55.0

[[columns]]
name = 'category'
type = 'string'
rows_count = 10000
missing = 0
modality_count = 25
min_length = 3
max_length = 45
avg_length = 12.5
top_value = "Electronics"
top_frequency = 2500
  

HTML Output Features

The HTML report includes:

  1. Overview Section:

    • Total rows and columns
    • Data types distribution
    • Missing values summary
  2. Numeric Columns Table:

    • Sortable by any statistic
    • Color-coded missing values
    • Distribution indicators
  3. Text Columns Table:

    • Cardinality indicators
    • Top values displayed
    • Length statistics
  4. Visualizations:

    • Missing values heatmap
    • Distribution histograms for numeric columns
    • Bar charts for top values in text columns

Best Practices

1. Always Explore Before Processing

Before writing any preprocessing logic:

  # Step 1: Skim the data
preprocess skim --data ./data.csv

# Step 2: Generate summary statistics
preprocess summary --data ./data.csv --html

# Step 3: Review the report
# Then design your preprocessing pipeline
  

2. Use Exclusions Wisely

Exclude columns that don’t provide valuable insights:

  preprocess summary --data ./data.csv \
  --exclude id \
  --exclude created_at \
  --exclude internal_notes
  

This makes your summary more focused and easier to read.

3. Save Both Formats

For comprehensive analysis, save both TOML and HTML:

  preprocess summary --data ./data.csv --output stats.toml
preprocess summary --data ./data.csv --html --output stats.html
  
  • Use TOML for programmatic analysis
  • Use HTML for presentations and sharing

4. Document Your Findings

Add notes to your summary files or include them in your project documentation:

  ## Data Quality Assessment

Based on summary statistics from `preprocess summary`:

- **Missing Values**: 5% of age values are missing
- **Outliers**: Salary values range from 20K to 5M (investigate upper range)
- **Cardinality**: Category column has 150 unique values (high cardinality)
- **Data Types**: All numeric columns detected correctly

**Actions Taken**:
- Impute missing ages with median
- Investigate salary outliers
- Consider dimensionality reduction for category column
  

Troubleshooting

Common Issues

Issue: No data source specified

  No data source specified. Please provide a data file or a prepfile.
  

Solution: Provide either --data or --prepfile:

  preprocess summary --data ./my_data.csv
# OR
preprocess summary --prepfile ./my_prep.toml
  

Issue: Encoding problems

  Error: Invalid UTF-8 encoding
  

Solution: Try different encodings:

  preprocess summary --data ./data.csv --encoding latin-1
  

Issue: Separator problems

  Error: Failed to parse CSV
  

Solution: Specify the correct separator:

  preprocess summary --data ./data.csv --sep ";"
  

Issue: HTML generation fails

  Error: Failed to generate HTML
  

Solution: Ensure you have write permissions in the directory and try again.

Performance Considerations

Large Datasets

For large datasets, summary generation can be resource-intensive:

  # For very large files, be patient
preprocess summary --data ./large_dataset.csv
  

Tips:

  • Exclude columns you don’t need with --exclude
  • Consider sampling your data for initial exploration
  • Use TOML format (faster than HTML)

Memory Usage

Summary generation loads the entire dataset into memory. For extremely large datasets:

  • Process in batches
  • Use the exclude flag to reduce memory footprint
  • Consider generating summaries for subsets of your data

Next Steps

After generating summary statistics:

  1. [Review the output]: Understand your data characteristics
  2. [Identify issues]: Look for missing values, outliers, data quality problems
  3. [Design preprocessing]: Use insights to inform your Prepfile configuration
  4. [Verify preprocessing]: Generate new summaries after preprocessing to confirm improvements
  5. Create Prepfile: Start building your preprocessing pipeline

Quick Reference Card

TaskCommand
Basic summary (TOML)preprocess summary --data ./data.csv
HTML reportpreprocess summary --data ./data.csv --html
Save to filepreprocess summary --data ./data.csv --output my_stats.toml
Exclude columnspreprocess summary --data ./data.csv --exclude col1 --exclude col2
Use Prepfilepreprocess summary --prepfile ./my_prep.toml
No browserpreprocess summary --data ./data.csv --html --no-browser
Custom separatorpreprocess summary --data ./data.csv --sep ";"
European decimalspreprocess summary --data ./data.csv --dsep ","

Summary Statistics Reference

Numeric Column Statistics

StatisticFormula/CalculationWhat It Tells You
minMinimum observed valueLower bound, potential outliers
maxMaximum observed valueUpper bound, potential outliers
meanSum of all values / countCentral tendency (affected by outliers)
medianMiddle valueCentral tendency (robust to outliers)
stdSquare root of varianceSpread/variability of data
q125th percentileLower quartile boundary
q375th percentileUpper quartile boundary
modeMost frequent valueMost common observation
missingCount of null/NA valuesData completeness

Text Column Statistics

StatisticCalculationWhat It Tells You
modality_countCount of unique valuesCardinality, potential for encoding
top_valueMost frequent stringDominant category
top_frequencyCount of top valueDistribution concentration
min_lengthShortest stringFormatting consistency
max_lengthLongest stringFormatting consistency, potential errors
avg_lengthAverage string lengthGeneral string characteristics
missingCount of null/NA valuesData completeness

Data Quality Indicators

IndicatorHow to IdentifyAction
High missing ratemissing > 10% of rows_countInvestigate, consider imputation or removal
Outliersmin or max far from q1/q3Investigate, consider winsorization or removal
High cardinalitymodality_count very highConsider grouping or encoding strategies
Zero variancemin = maxFeature provides no information, consider removal
Inconsistent lengthsLarge range between min_length and max_lengthInvestigate formatting, potential data entry errors

Last updated 18 Sep 2026, 11:59 +0200 . history