No description
Find a file
2026-09-01 12:55:55 +02:00
data add examples 2026-09-01 12:41:21 +02:00
examples add examples 2026-09-01 12:41:21 +02:00
src/dp_tests add examples 2026-09-01 12:41:21 +02:00
.gitignore init 2026-07-23 13:32:27 +02:00
main.py init 2026-07-23 13:32:27 +02:00
pyproject.toml Update package name to dp-tests and add build system 2026-09-01 12:55:55 +02:00
README.md add install instructions 2026-09-01 12:49:10 +02:00
uv.lock rework as lib module 2026-09-01 12:07:02 +02:00

DP Tests

A lightweight, extensible benchmarking suite for Differential Privacy (DP) tabular synthesizers and privatizers built with a BYODP (Bring Your Own DP) philosophy.

You handle the privacy algorithm; dp_tests handles all the data loading, encoding, train/test splits, epsilon sweeps, dataset size sweeps, baseline bounds, fidelity metrics, and downstream ML utility evaluations.

Installation

Add dp_tests to your project with uv:

uv add git+ssh://git@git.leon-adsk.com:222/leon-adsk/dp_tests.git

Quickstart

Run an evaluation with the dp_synth synthesizer on the Adult Income dataset:

import dp_tests as dpt

# Initialize experiment from a dataset schema
exp = dpt.Experiment(
    schema="data/schemata/adult_income.yaml",
    dp_method="dp_synth",
    dp_params={"k_max": 2},
)

# Run baseline models on real data (upper, lower, and standard bounds)
baselines = exp.run_baselines()
print(baselines[["model", "bound", "accuracy", "roc_auc", "f1_score"]])

# Run an epsilon sweep across privacy budgets
results = exp.run_epsilon_sweep(
    epsilons=[0.1, 1.0, 10.0],
    iterations=3,
)
print(results[["model", "epsilon", "fidelity", "accuracy", "roc_auc", "runtime"]])

Core Features

  • Bring Your Own DP Method (BYODP): Plug in any DP synthesizer, local perturbation mechanism, or generative model via a simple callable interface.
  • Custom Downstream ML Models: Benchmark synthetic data utility on your choice of models (Scikit-Learn, XGBoost, LightGBM, custom neural nets, or standard defaults).
  • Automated Sweeps:
    • Epsilon Sweeps: Test utility/fidelity across varying privacy budgets \epsilon \in [\epsilon_1, \dots, \epsilon_n].
    • Dataset Size Sweeps: Test robustness across subsampled dataset sizes (e.g. 10%, 25%, 50%, 100%).
  • Multi-Bound Baselines:
    • Upper Bound: Performance using all available features.
    • Baseline: Performance on selected data types (NUMERIC, ORDINAL, CATEGORICAL).
    • Lower Bound: Performance on inverse/excluded features.
  • Spurious Correlation & Poisoning Tests: Inject synthetic spurious artifacts to measure whether DP mechanisms suppress or amplify dataset biases.
  • Target Swapping: Evaluate synthetic data fidelity when predicting non-target columns.
  • Fidelity & Utility Metrics: Computes Total Variation Distance (TVD) marginal fidelity alongside classification Accuracy, Macro F1, and ROC AUC.

Comprehensive Example: Custom DP, Custom Models & Spurious Correlations

Here is a full example demonstrating how to combine:

  1. A custom Local DP mechanism (Laplace noise).
  2. Custom ML evaluation models (LogisticRegression, KNeighborsClassifier, and GradientBoosting).
  3. Spurious correlation injection to test bias mitigation under differential privacy.
from pathlib import Path
import numpy as np
import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.neighbors import KNeighborsClassifier
from sklearn.ensemble import GradientBoostingClassifier

import dp_tests as dpt
from dp_tests.data.dataset import DatasetSchema


# Define a Custom DP Method (BYODP)
def my_laplace_privatizer(
    df: pd.DataFrame,
    epsilon: float,
    schema: DatasetSchema,
    seed: int | None = None,
    **kwargs,
) -> pd.DataFrame:
    """Adds calibrated Laplace noise to numeric columns based on schema bounds."""
    df_priv = df.copy()
    rng = np.random.default_rng(seed)

    for col in schema.num:
        if col in df_priv.columns and col in schema.bounds:
            min_val, max_val = schema.bounds[col]
            sensitivity = max_val - min_val
            scale = sensitivity / max(epsilon, 1e-6)

            noise = rng.laplace(0.0, scale, size=len(df_priv))
            df_priv[col] = np.clip(df_priv[col] + noise, min_val, max_val)

    return df_priv


# Define Custom Downstream ML Models
custom_eval_models = {
    "LogisticRegression": LogisticRegression(max_iter=1000, random_state=42),
    "KNN (k=5)": KNeighborsClassifier(n_neighbors=5),
    "GradientBoosting": GradientBoostingClassifier(n_estimators=50, random_state=42),
}


# Configure the Experiment
exp = dpt.Experiment(
    schema="data/schemata/adult_income.yaml",
    dp_method=my_laplace_privatizer,
    models=custom_eval_models,
    
    # Inject spurious correlation into a feature to test bias robustness:
    spurious_col="relationship",
    spurious_strength=0.85,
    
    results_path=Path("data/results/"),
    verbosity="milestones",
)


# Run Baselines (including poisoned baseline)
baseline_df = exp.run_baselines()
print("=== Baseline Results ===")
print(baseline_df[["model", "bound", "accuracy", "roc_auc", "f1_score"]])


# Run Privacy Budget Sweep
results_df = exp.run_epsilon_sweep(
    epsilons=[0.1, 1.0, 10.0, 50.0],
    iterations=3,
)

print("\n=== Epsilon Sweep Results ===")
print(results_df[["model", "epsilon", "fidelity", "accuracy", "roc_auc", "runtime"]])

Dataset Schemata

Datasets are configured using simple YAML schema files placed in data/schemata/:

target: "income"
cat:
  - "workclass"
  - "education"
  - "marital-status"
  - "occupation"
  - "relationship"
  - "race"
  - "sex"
  - "native-country"
num:
  - "age"
  - "fnlwgt"
  - "capital-gain"
  - "capital-loss"
  - "hours-per-week"
ordinal:
  - "education-num"
bounds:
  age: [17, 90]
  fnlwgt: [12285, 1484705]
  capital-gain: [0, 99999]
  capital-loss: [0, 4356]
  hours-per-week: [1, 99]
  education-num: [1, 16]