Idle databases spun down after 24h on Datomic Cloud cause "Loading database" on the next transaction

Setup

  • Datomic Cloud 1254-9433, production topology, us-west-2.
  • Primary compute group api2: 2 x i3.2xlarge. Query group web: 1-3 x i3.xlarge.
  • About 14 databases. PreloadDb is empty on both groups.
  • Ions use the in-process client (datomic.cloud.client.local.Connection).

Symptom

Every week or two, a d/transact against a rarely used database fails with:

{:cognitect.anomalies/category :cognitect.anomalies/unavailable
 :cognitect.anomalies/message  "Loading database"}

This happened on 1254-9433 as well as on our previous release. Correlating with the system’s CloudWatch logs shows that each node unloads a database after 24 hours without use. The first transaction afterwards is rejected while that node reloads the database.

Evidence: 2026-10-01 (UTC)

DbId 3dcee24b-c74a-4cea-bc8e-186ea2a4cb16 is a small database (basis-t about 3500) that only gets written when users use one feature.

  1. At 2026-09-30 20:24-20:28, all four nodes started after an ion deploy and loaded the DB (LogCatchup / SpinupDb). After that, there were no TxSucceeded or other log events for this DbId on any node for 24 hours.

  2. At 20:29:53.786, on primary i-00f626b67bb6f6a7e, GroupViewDbsChanged showed "3dcee24b-...": "Wed Sep 30 20:29:52 UTC 2026" in Old. The DB was absent from New. Then:

    SpindownDbIds {"DatomicCloudUpdateSystemCacheDbids": "#{\"3dcee24b-c74a-4cea-bc8e-186ea2a4cb16\"}"}
    SpindownDb    {"DbId": "3dcee24b-...", "BasisT": 3518, "IndexBasisT": 5}
    

    The other primary, i-093f376c70cdb30e0, did the same at 20:30:35.121. Other DBs were spun down the same way, each about 24h after the timestamp shown in GroupViewDbsChanged. For example, f586a532-... had last use 09-30 20:29:54 and was spun down 10-01 20:30:53.

  3. The first transaction afterwards, on i-00f626b67bb6f6a7e:

    20:34:55.369 ClientSPIAnomaly {"DatomicClientSpiErrorResponse": {"Status": 200, "Body":
                   {"CognitectAnomaliesCategory": "CognitectAnomaliesUnavailable",
                    "CognitectAnomaliesMessage": "Loading database"}}}
    20:34:55.376 IndexLoaded {... "IndexRootId": "b56ea0e1-..."}
    20:34:55.377 LogCatchupRequest {"NextT": 6, "DbId": "3dcee24b-..."}
    20:34:55.633 LogCatchup {"NextTAfter": 3519, "Msec": 256, "Bytes": 7871334}
    20:34:55.634 SpinupDb {"DbId": "3dcee24b-...", "BasisT": 3518}
    20:35:02.380 TxSucceeded {"DbId": "3dcee24b-...", "T": 3519}
    

    The anomaly was returned about 265 ms before the database became available on that node.

  4. The second primary reloaded the DB at 20:35:53 without any failed request.

Evidence: 2026-09-22, previous Datomic release

The same pattern occurred, with one difference: the query-group nodes reloaded the DB for reads 11 s before the write still failed on a primary.

03:53:21.252 web-i-0345d97ad2aa78a80   SpinupDb 3dcee24b-...
03:53:24.587 web-i-0e3e8156a3e3a313c   SpinupDb 3dcee24b-...
03:53:32.797 api2-i-0c761bf4836838480  ClientSPIAnomaly "Loading database"
03:53:35.378 api2-i-0c761bf4836838480  LogCatchup 3dcee24b-... Msec=2573
03:53:35.379 api2-i-0c761bf4836838480  SpinupDb 3dcee24b-...

The primaries had spun the DB down on 2026-09-19 at 09:09-09:10.

We also see last-use timestamps on a primary advance with no transaction for that DbId on any node within +/-2 s. For example, e10ff883-... at 2026-10-01 17:12:37 on i-00f626b67bb6f6a7e. That suggests reads also count as use.

Questions

  1. Is the 24-hour idle spindown intended behaviour? Is the threshold documented or configurable?
  2. What counts as “use” for the last-use timestamp in GroupViewDbsChanged? Do reads count, does d/connect / d/db alone count, and is it tracked per node?
  3. Does PreloadDb exempt that database from idle spindown, or does it only load the database at instance start? It accepts a single name. Is there a supported way to keep several databases loaded?
  4. Why is a transaction to a spun-down database rejected immediately with unavailable, instead of waiting for the load (about 250 ms here)?
  5. Is it guaranteed that a transaction rejected with “Loading database” was not applied? We want to retry these with backoff, as the troubleshooting page recommends for connect, and need to know that retrying transact cannot apply a transaction twice.

Thanks!

(It’s ongoing issue for years by now and recently we lost OAuth access and refresh tokens due to it. :frowning: )

  1. Yes, it is intended, but not publicly documented or configurable. It is currently 24 hours. To reveal a reason for why this exists is that it is closely tied to garbage collection and java runtime maintenance in addition to preserving used/rarely used performance gains for large multi tenant systems.

  2. This is a generic hook that is on d/db, query, and transact paths. It’s not transact only. It is tracked per node and you will note in your logs you provided that the log’s you captured are per instance. So you could/should implement a d/db conn/no-op read or write to ensure the DB is loaded periodically ahead of the expected transaction.

  3. No PreloadDB will not exempt the DB from spindown. The only supported method to keep a DB loaded is periodic reads or a no-op write on the db in question. This is an area we can work on a feature to pin a DB as always loaded, but I suspect there are tradeoffs and I would need to consider this feature with development.

  4. We should more gracefully handle this scenario, but the current design is to reject once the DB is not loaded immediately without blocking. Larger Databases can take longer times to load and we did not want to block the hot path.

  5. This operation is safe to retry. If you received an unavailable/loading database error the transaction will not be applied. We follow the general anomaly patterns where possible GitHub - cognitect-labs/anomalies · GitHub.

It’s ongoing issue for years by now and recently we lost OAuth access and refresh tokens due to it. :frowning:

Can you help me understand what happened here? Was this because the initial transact failed and you did not notice the failure in time or retry? I just would like to understand how this happened.

ADDITIONAL INFORMATION:

{ $.Msg = "SpindownDb" || $.Msg = "SpinupDb" || $.Msg = "SpinupWaitCompleted" || $.Msg = "SpinupWaitAbandonded" }

These log messages events allow you to view the cycle of spin up and spin downs on your nodes via CloudWatch Logging.

While it is possible for a db set in preLoadDb to cycle down (no use for 24 hours on a long lived instance) it is less likely as all startup scenarios (ion deploy, restart) will ensure that the databases is loaded initially.

The points you have raised are worth our team exploring for features to improve this behavior. I am discussing with development. I may have more to post later.

@onetom I want to add that if you need a rock solid guarantee on the retry for the transact you should put a cas or mark it as a unique value to ensure that you are not misusing the retry. Happy to talk about what that could look like in your specific domain.

Also after talking with development we believe the query path to “loading the db” might get served by another node. If you issue an empty transaction though you will ensure that the db begins loading. You could submit empty transactions once an hour or on a schedule to ensure you avoid the spindown.