Ratios (debt / income), differences (days since last purchase), aggregates per entity (a customer's average order value), interactions (price × quantity) and flags (is_weekend) often add more than any algorithm change.
Domain knowledge tells you which ones matter. Ask: what would a human expert look at to make this decision?
Measure every feature honestly: keep a fixed CV setup and compare scores with and without it. Aggregates per entity must be computed without peeking at the target row or the future.
Going deeper
Time-aware aggregates (a customer's spend in the previous 30 days, computed as of each prediction date) are the most powerful features in transactional data, and the most leakage-prone. Build them with explicit 'as of' timestamps.
Automated feature tools (Featuretools) generate candidates by stacking aggregations over relationships, but domain-driven features still tend to win.
Common pitfalls
- Aggregates that include the current row's label (leakage).
- Adding dozens of features without measuring each one's contribution.