unfazed
WIUT Hackathon / 2026EXPLORATORY DATA ANALYSIS

The team
B7F53A18

Behind every alert, a transaction history.
We explored it to rank the likelihood of escalation.

Explore the findings

Synthetic data.
Real experiments.
About 5 minutes to explore.

10,015,238
historical transactions
20,000
unique alerts
14,000 train / 6,000 test
180 days
maximum lookback
01The starting point

Most alerts end
without escalation.

One alert, many transactions, one outcome. Our unit of prediction is the alert—not the individual transaction.

The target is imbalanced

14,000 training alerts · Jan 2025–Dec 2026

01.1
17.18%

were escalated

2,405 escalated
11,595 dismissed
EscalatedDismissed

Predicting “dismissed” for everyone would miss every escalation. We care about ranking, not a majority-class shortcut.

What moves through the account?

Share of transaction rows, not alert counts

01.2

Cards and bank transfers dominate. International transactions are rare, making their per-alert summaries less stable.

0 overlapbetween train and test IDs
0 future rowsafter the associated alert date
0 raw nullsin the columns examined
02Inside the history

Similar behavior.
Subtle differences.

No single field cleanly separates the outcomes. Explore what changes—and what mostly overlaps.

DismissedEscalated
Move or tap across the chart to inspect a bin.

More transactions. Not a different universe.

Escalated alerts had slightly more activity on average, but the distributions overlap heavily.

494Dismissed mean
524Escalated mean

Transactions per alert.

Each class is normalized separately. Lines join histogram-bin centers; they are not fitted probability curves.

Approaching the alert

TRAIN ONLY

Daily transaction activity relative to the signal date

DismissedEscalated
A sharp day-before-alert spike appears in both groups. Yet the average days 0–7 share barely separates outcomes: 10.50% dismissed versus 10.29% escalated.

Per-alert means; each point is one calendar day. Day −1: 40.10 transactions per dismissed alert, 41.66 per escalated alert. Days 0–7 include eight calendar dates.

A comparable hidden set

TRAIN ↔ TEST

Cumulative distribution of transactions per alert

TrainTest
The two activity curves nearly overlap. We found no obvious shift in the aggregate features examined—not a guarantee that every pattern matches.
461 / 460.5 median transactions · train / test
03From rows to a ranking

Turn a history
into a description.

Millions of transactions, but only 14,000 labeled examples. We first needed an alert-level representation.

[ 01 / AGGREGATE ]

Keep the history together.

Join each alert to its transactions. Summarize the history before the alert date, rather than treating every transaction as an independent training example.

DuckDB + Parquetone row per alert
[ 02 / DESCRIBE ]

Look within the categories.

Measure size distributions inside each transaction type and direction. Add activity, variability, recent changes and sequence summaries. Then test which groups actually help.

size quantilesdirection × typebehavioral ratios
[ 03 / VALIDATE ]

Learn. Hold out. Repeat.

Fit on four folds and score the fifth. Imputation and scaling learn only from each training fold. Repeat across three split seeds, then fit the selected model on all training alerts.

5-fold stratified CVfold-local preprocessing
THE REPRESENTATION01 / 03
Many transactions → one alert

The size field is a standardized indicator—not currency. Negative values are valid; sums of this indicator are not financial turnover.

04What actually helped

More features
weren’t always better.

We removed feature groups to see what the model depended on. Category-specific size distributions made the biggest difference.

Remove a group. Watch the score.

FEATURE ABLATION

Compact logistic model · C = 0.01 · same five folds, split seed 67

OOF ROC-AUC · zoomed axis starts at 0.54. Baseline: 0.63022.

0.56437
-0.06586 AUC · 193 features remain

Remove category-conditioned size summaries and compare the ranking.

Groups overlap: removing composition also removes some recent-share changes. These are group diagnostics, not isolated causal effects.

The focused alternative: 57 size-focused features + regularized logistic regression. It reached 0.64100 AUC on split 67. LightGBM adds complementary information in a blend.

Explore the model experiments 19 configurations
Representation / modelFeaturesOOF AUC

Same five folds · split seed 67.

05The final solution

A small blend.
A consistent improvement.

Size-focused logistic regression does most of the work. A smaller LightGBM contribution improves the ranking.

0.64512
out-of-fold ROC-AUC · not accuracy
+0.01226 over the original blend
FINAL PROBABILITY MIX
Logistic regression57 size-focused features
LightGBM439 enriched features

A weighted average of probabilities. Both components are finally fitted on all 14,000 labeled alerts.

CHANGE SPLIT SEED

Ranking across thresholds

ROC CURVE
Final blendLogisticLightGBM
Higher is better. The dashed diagonal is chance. Curves use sampled points for display; AUC uses every out-of-fold prediction.

The same folds. A better ranking.

OOF AUC

Bars show improvement above chance (AUC 0.50); the displayed range ends at 0.67.

5 folds × 3 split seeds

AUC ≈ 0.645 means an escalated alert ranks above a dismissed one in about 64.5% of random pairs. It is not classification accuracy.

Validation details & model settings
Split seedOriginal blendFinal ensembleGain

The reference is the original 50/50 logistic / histogram-boosting configuration, recomputed with model seed 67 on matching folds.

A paired stability check

On split 31415, the fixed 75/25 ensemble gained 0.01069 AUC. A 500-resample paired bootstrap gave a 95% interval of [0.00448, 0.01634].

This interval is conditional on fitted predictions and does not account for the full model search. Repeated splits reuse the same labels; they are stability checks, not independent holdout datasets.

Reproducible, without rewriting labels

Logistic: median imputation, missingness indicators, scaling, C=0.1. LightGBM: 350 trees, 7 leaves, learning rate 0.025, L2=10. Model and bootstrap seed 67. Timestamp ties use a canonical ordering.

Scores and curves use the executed notebook with split seeds 67, 2026 and 31415. Small numerical differences are possible on reruns.

WHAT WE TAKE AWAY

Look inside the distribution.

01 / CONTEXT MATTERS

Size within a category.

Direction- and type-specific size distributions were more useful than a single overall description of activity.

02 / TEST THE INTUITION

Complexity has to earn its place.

Adding recent changes and more behavioral summaries did not automatically help. Feature removal improved the linear model.

03 / MEASURE THE GAIN

Small, but repeatable.

The ensemble gained 0.011–0.012 AUC over the original baseline across the three fold assignments.