[Bug] wing.db-journal left on unclean shutdown causes "database is locked" on restart; please enable WAL mode + busy_timeout
Environment
daeuniverse/dae-wing running as the embedded GraphQL backend inside daeuniverse/daed
wing.db is the SQLite database for dae-wing's domain routing and subscription state
- iStoreOS / OpenWrt 6.6.119 on aarch64
- 4 GB RAM + 2 GB swap
Summary
This is a more specific follow-up to issue #489 (which fixed a nil-pointer + SQLITE_BUSY in 0.7.0). We still occasionally see wing.db-journal left behind after unclean shutdown, leading to database is locked (SQLITE_BUSY (5)) on the next start, with the same ultimate symptom: daed loads but cannot serve any GraphQL query, so subscription refresh / node updates silently fail.
The root cause appears to be a long-lived write transaction in dae-wing holding the journal mode's exclusive lock, combined with no busy_timeout configured on the SQLite connection. Issue #489 closed the immediate panic, but the underlying connection setup is unchanged.
Evidence
- 2026-08-25 overnight-watch alerts log captured a
wing.db-journal left in /etc/daed/ after a forced kill -9 of the daed process, followed by an SQLITE_BUSY on the next start. The journal was not cleaned up by the wrapper; only a full reboot cleared it.
- GraphQL subscription refresh failed silently for ~40 minutes on 2026-08-25 11:00–11:40 with no log line beyond the original panic; recovery was via reboot.
- The same
wing.db-journal reappeared twice in the 2026-08-25 window across reload attempts.
Steps to reproduce
- Run daed for a few days under steady traffic.
- Force a sudden stop (
kill -9 on the daed process, or a hard power cycle on the router).
- The next start produces:
panic occurred: runtime error: invalid memory address or nil pointer dereference
database is locked (5) (SQLITE_BUSY)
ls /etc/daed/wing.db* shows a wing.db-journal file lingering.
- The journal is not cleaned up by
daed-guard or the procd init script; only a full reboot clears it.
Expected
- The connection should set
PRAGMA busy_timeout = 5000 so transient contention waits instead of failing immediately.
- The connection should set
PRAGMA journal_mode = WAL so readers do not block writers (and to avoid journal-mode lock contention that triggers the SQLITE_BUSY in the first place).
- On unclean shutdown, dae-wing should run a
PRAGMA wal_checkpoint(TRUNCATE) so -wal and -shm do not leak.
Suggested fix
In the SQLite connection initialisation (wherever dae-wing opens wing.db):
db.Exec("PRAGMA busy_timeout = 5000")
db.Exec("PRAGMA journal_mode = WAL")
And on shutdown (in the dae-wing Run() exit path):
db.Exec("PRAGMA wal_checkpoint(TRUNCATE)")
Impact
Anyone running daed as their long-lived transparent proxy. The lock shows up only after unclean shutdowns (OOM, power loss, watchdog), so it is by definition the most operationally important time for it to work.
[Bug] wing.db-journal left on unclean shutdown causes "database is locked" on restart; please enable WAL mode + busy_timeout
Environment
daeuniverse/dae-wingrunning as the embedded GraphQL backend insidedaeuniverse/daedwing.dbis the SQLite database for dae-wing's domain routing and subscription stateSummary
This is a more specific follow-up to issue #489 (which fixed a nil-pointer + SQLITE_BUSY in 0.7.0). We still occasionally see
wing.db-journalleft behind after unclean shutdown, leading todatabase is locked(SQLITE_BUSY (5)) on the next start, with the same ultimate symptom: daed loads but cannot serve any GraphQL query, so subscription refresh / node updates silently fail.The root cause appears to be a long-lived write transaction in dae-wing holding the journal mode's exclusive lock, combined with no
busy_timeoutconfigured on the SQLite connection. Issue #489 closed the immediate panic, but the underlying connection setup is unchanged.Evidence
wing.db-journalleft in/etc/daed/after a forcedkill -9of the daed process, followed by anSQLITE_BUSYon the next start. The journal was not cleaned up by the wrapper; only a full reboot cleared it.wing.db-journalreappeared twice in the 2026-08-25 window across reload attempts.Steps to reproduce
kill -9on the daed process, or a hard power cycle on the router).ls /etc/daed/wing.db*shows awing.db-journalfile lingering.daed-guardor the procd init script; only a full reboot clears it.Expected
PRAGMA busy_timeout = 5000so transient contention waits instead of failing immediately.PRAGMA journal_mode = WALso readers do not block writers (and to avoid journal-mode lock contention that triggers the SQLITE_BUSY in the first place).PRAGMA wal_checkpoint(TRUNCATE)so-waland-shmdo not leak.Suggested fix
In the SQLite connection initialisation (wherever dae-wing opens
wing.db):And on shutdown (in the dae-wing
Run()exit path):Impact
Anyone running daed as their long-lived transparent proxy. The lock shows up only after unclean shutdowns (OOM, power loss, watchdog), so it is by definition the most operationally important time for it to work.