v2.1.5: memory climbs to the cgroup limit and gets OOM-killed every 1–2 days (Docker, ~430k files)

Hi,

I’m running Syncthing in Docker on a Synology NAS and it keeps hitting its memory limit. The kernel OOM killer takes out the syncthing process, the built-in monitor restarts it, and Docker never notices (container stays Up, RestartCount=0).

Environment

  • Syncthing v2.1.5 “Hafnium Hornet” (go1.27.1 linux-amd64), official image syncthing/syncthing:latest
  • Host: Synology DSM 7.4.1 (build 90080), kernel 4.4.302+, Ryzen R1600 (2c/4t), 32 GB RAM
  • Container: mem_limit / memswap_limit = 2 GiB (now 4 GiB), GOMEMLIMIT=1600MiB (now 3200MiB)
  • One remote device (macOS, v2.1.5), LAN only
  • index-v2 database: ~956 MB

Folders (7, all send-receive except one)

Folder Files Size
A 271,824 243 GB
B 90,436 98 GB
C 63,213 9.4 GB
D 5,188 5.7 GB
E 654 0.8 GB
F 246 2.6 GB
G 99 5.7 MB

~431k files / ~359 GB total. fs watcher enabled on all folders; rescan interval 120 s on two of them (NFS-backed, no inotify), 3600 s on the others.

What happens

Kernel log each time:

syncthing invoked oom-killer: gfp_mask=0x2400040, order=0, oom_score_adj=0
mem_cgroup_oom_synchronize+0x2c2/0x2f0
Killed process 4189 (syncthing) total-vm:3812372kB, anon-rss:2073532kB, file-rss:0kB

Syncthing log right after:

INF Syncthing exited (error="signal: killed" log.pkg=main)

Nothing unusual is logged in the minutes before any of the kills. Anonymous RSS sits right at the limit every time.

The memory is not in the Go heap. Snapshot taken 18 h after the last restart, with the process close to the 4 GiB limit:

VmRSS            3,732,192 kB     (VmSwap 73,480 kB)
cgroup total_rss 3,813,568,512    total_cache 385,630,208
runtime Sys        227,264,824
HeapInuse          110,272,512    HeapIdle 97,280,000   HeapReleased 94,445,568
goroutines         142

So roughly 3.5 GB of anonymous memory lives outside the Go runtime, which is why GOMEMLIMIT has no effect. My guess is the SQLite index (the binary is static, so the SQLite allocator isn’t the Go heap), but I haven’t confirmed it.

/proc/<pid>/smaps of the syncthing process, 3 h after the last restart (RSS 1.96 GiB, HeapInuse 66 MB), anonymous mappings only:

mapping size    count   RSS
>= 64 MB            4   108 MB    <- the Go arenas
4-64 MB           108   412 MB
1-4 MB            602   843 MB
< 1 MB          3,981   633 MB
file-backed total        16 MB

~4,700 small anonymous mappings hold ~1.9 GB while the Go arenas hold ~108 MB. That looks like an allocator that mmaps on its own, outside the Go runtime. If the release binary uses the pure-Go SQLite (modernc), is its libc allocator the one growing here?

Kills

At 2 GiB, 10 kills in 9 days (UTC):

2026-09-19 15:40   2026-09-23 03:35   2026-09-23 14:40   2026-09-23 20:26
2026-09-24 06:31   2026-09-25 09:49   2026-09-25 12:53   2026-09-28 01:21
2026-09-28 09:53   2026-09-28 12:27

At 4 GiB (since 2026-09-28 13:31 UTC), 2 kills so far:

2026-09-28 17:43   2026-09-29 12:24

The second one came 15 minutes after the snapshot below (RSS 3.73 of 4 GiB). The cgroup’s memory.max_usage_in_bytes reads 4,298,772,480 — right at the limit.

Questions

  1. Is ~3.5 GB of non-Go memory expected for ~430k files with the v2 SQLite index (total DB size ~1 GB)? Is there a setting that bounds the SQLite cache / memory?
  2. If not, what would help you most? Attached: the snapshot above plus a heap profile from STPROFILER (/debug/pprof/heap) taken at the same moment, and the full smaps capture.

Thanks!

This seems unexpected. I have about half the amount of files, though more data, but my Syncthing uses 22 MiB RSS. You’re probably right about the non-Go allocations being SQLite. SQLite cache memory is capped at 2 MiB per connection, with four persistent connections and a maximum under busy times of 16 or 4*NumCPU, whichever is higher. You’d have to have a lot of CPUs, and reason for Syncthing to use a lot of database connections, for that to be cache memory though…

If the release binary uses the pure-Go SQLite (modernc), is its libc allocator the one growing here?

I don’t know, but our binaries (and hence the container) use the embedded C SQLite.

Thanks — that corrects my guess: the binary does contain go-sqlite3, so it’s the embedded C SQLite, statically linked.

Some numbers against your cache math: v2.1 here keeps 8 databases (main + one per folder), 23 open DB handles right now, 15 threads, NumCPU=4. Even at the burst maximum that’s ~256 MiB of cache, nowhere near 3.5 GB.

New data point: a third kill at 4 GiB this morning. Growth isn’t linear; it comes in bursts: RSS went 3186 → 3729 MiB in under 3 minutes, then the process was killed 20 min later.

Something I hadn’t noticed: the remote device reaches this container over two paths (outbound to its LAN IP, classified as WAN because of Docker bridge networking, and inbound via the bridge gateway), and they keep replacing each other. That’s 112 connections in 42 h, with “Abandoning old index handler in favour of new connection” among them. Could repeated index-handler restarts allocate in SQLite/cgo without releasing it? I’ll try to remove the dual path and report whether the bursts stop.