retention · raw request logs
Keep most raw request logs for 45 days
Ordinary raw logs fall from 400 days to 45; 5xx logs keep 120 days. Aggregates stay at 400 days. The change returns about 18 TB of hot storage and £3,000 a month after the 5xx carve-out.
The aggregates and audit store do not change
Per-minute status, latency, route, and region aggregates keep 400 days. Audit records remain in their separate seven-year store.
Most raw-log reads are within six weeks
Across eighteen months of query history, 90% of reads touch logs less than a week old and 95% stay within 42 days. Query history cannot reveal work that retention prevented, so I checked two years of incident write-ups as well.
The 45-day limit rounds the observed six-week boundary up by three days.
What older incidents needed
Two years of incident write-ups name a log older than six weeks exactly twice. Aggregates supplied the rate both times. In March, the raw request id was also needed to identify four affected accounts rather than only count them.
Keep 5xx request logs for 120 days. They are 4% of raw-log volume and cover both incidents; other raw logs expire at 45 days.
Two-week rollout
- Week one: run expiry in report-only mode and publish deletions by team.
- Also in week one: find dashboard queries that read raw logs directly.
- Week two: ship the 5xx carve-out before deleting anything.
- Week two: expire the oldest data over four days and watch query failures.