← Back to Notes

The Rough Edges of Upgrading a Self-Hosted AI Agent

A firsthand upgrade exposed how database bloat, live state, schema changes, and indexing can turn a package update into a full maintenance event.

Self-hosted AI agent components passing through backup, upgrade, and verification stages during a maintenance window.

Self-hosting an AI agent can make the system feel unusually personal. The agent has local files, accumulated memory, tools, scheduled work, and a history that grows alongside the installation. That continuity is the point.

It is also what makes an upgrade more complicated than replacing an application package.

We learned this during a September upgrade of our OpenClaw installation. The core update eventually landed, but the path exposed maintenance assumptions that had been easy to miss while the system was healthy. A large local logging database made validation slow. Live processes kept writing while diagnostic tools tried to inspect state. A memory configuration that had worked before no longer matched the new schema. Automatic indexing then made it hard to tell whether a test was observing the current system or one of several intermediate states.

None of those problems, by itself, meant the software was fundamentally broken. Together, they turned an ordinary upgrade into a real maintenance event.

The Package Was Only One Layer

The first mistake in thinking about a self-hosted agent upgrade is to treat it like a normal desktop application update.

An agent installation is closer to a small service stack. The package is one layer, but it may sit on top of persistent databases, model files, vector indexes, messaging connections, scheduled jobs, plugins, tool credentials, and long-running worker processes. Some of those components are rebuilt automatically. Others carry years of local history. Several may be active while the updater is trying to inspect or migrate them.

That means a successful package install does not prove the system is operationally healthy. The new executable may be present while a plugin migration is pending, a database is under contention, or retrieval is returning nothing.

Our upgrade made that distinction visible. The installed version advanced, yet post-install verification still encountered state inspection and migration trouble. The right question was no longer “Did the package update?” It was “Can the whole system still perform the work we depend on?”

A 22.8 GB Database That Mostly Contained Empty Space

The clearest storage problem was Main’s Codex log database. On disk, logs_2.sqlite occupied 22,830,137,344 bytes, or about 22.8 GB.

That sounded like an enormous amount of live history. It was not.

SQLite reported 5,573,764 pages, with 5,448,472 pages on the freelist. In other words, old records had been deleted logically, but the empty pages were still allocated inside the file. The database was healthy; the file simply had not returned most of that space to the filesystem.

This matters because tools generally pay some cost for the physical shape of a database, not only for the rows we think of as current. Backup attempts, integrity checks, schema inspection, and broad filesystem searches all became slower or less predictable around a file of that size. A first attempt at an inspection snapshot was too slow to be useful. A first cleanup attempt also recovered almost nothing because one incremental-vacuum call removed only one page in this environment.

The eventual maintenance procedure was deliberately boring: stop the gateway and Codex workers, verify a backup, run a full SQLite VACUUM, check integrity and row counts, and restart the service. The database fell from 22,830,137,344 bytes to 439,513,088 bytes—about 420 MB. The operation recovered 22,390,624,256 bytes, and SQLite’s quick check returned ok.

The lesson is not that every large SQLite file should be vacuumed. A full vacuum needs free space, exclusive-enough access, downtime planning, and a verified backup. The lesson is that retention and compaction are different jobs. Deleting old rows can satisfy a retention rule while leaving the physical file almost unchanged.

It is also important not to merge unrelated databases into one story. The oversized Codex log database did not store the memory embeddings that later failed. It was a separate source of operational friction that slowed the broader maintenance work.

Live State Made Every Observation Less Reliable

The upgrade’s safeguards were doing real work: rehearsing migrations, validating configuration, resolving plugins, launching a gateway canary, running doctor checks, and inspecting state. Several of those checks completed. Others ran into slow SQLite transactions, timeouts, or state that would not stabilize long enough to inspect.

The system was not idle while this happened. Gateway activity, Codex workers, scheduled automations, plugin migrations, and indexing could all write state. A check could begin against one generation of the process and finish after another had started. A restart could fix one issue while invalidating the observation that motivated it.

This is where repeated probing becomes counterproductive. Each extra status command seems harmless, but on a busy installation it can add work precisely when the system needs a quiet period to finish migrations and settle. We saw stale observations after restarts and had to rerun verification instead of trusting an earlier result.

A maintenance window is therefore more than permission for downtime. It is a controlled reduction in concurrency. Stop or stagger cron-heavy work. Let the updater finish. Avoid launching parallel diagnostics until migrations have completed. If a database must be compacted or rebuilt, make sure the processes that normally hold it open are actually stopped.

The Memory Failure Was Not One Failure

After the upgrade, semantic memory and session retrieval did not behave normally. It was tempting to name a single cause quickly: the embedding provider, the local model runtime, the index, or the large SQLite file.

The evidence pointed to a layered incident instead.

A legacy memory configuration key was rejected by the new schema. At the same time, automatic reindexing and continuing state writes created contention and timeouts. Some early conclusions about the embedding provider were premature because tests were being run while the system was still changing.

The final working configuration used OpenClaw’s local provider with both memory and session sources and the existing local EmbeddingGemma model. We verified both retrieval paths with real queries. A temporary compatibility service created during troubleshooting was removed, and retrieval was tested again so it could not be mistaken for a required part of the installation.

That final verification mattered more than a green service status. A process can be running while memory search is disabled, stale, or returning the wrong corpus. If retrieval is a feature you depend on, the acceptance test has to be retrieval itself.

What the Safeguards Could and Could Not Do

It would be unfair to describe the updater as having no protections. The logs show migration rehearsal, configuration validation, plugin resolution, a gateway canary, doctor checks, repair guidance, and rollback information. Those mechanisms identified real problems and preserved state instead of blindly forcing a result.

But safeguards have operating assumptions. They need databases that can be inspected within a reasonable bound. They need configuration that can be migrated unambiguously. They need enough quiescence to distinguish a slow check from a moving target. A long-running customized installation can violate several of those assumptions at once.

This is the rough edge of self-hosting: the operator owns the unusual history of the system. Upstream can provide migration logic and diagnostics, but it cannot know that one local log database has millions of free pages, that dozens of scheduled jobs are active, or that a compatibility service created during troubleshooting should not survive the incident.

The answer is not to avoid upgrades. It is to treat them as application, data, model, index, and orchestration migrations together.

A Practical Upgrade Checklist

Our next upgrade should begin before the updater runs:

  • Inspect persistent stores. Record database file sizes, page counts, freelist counts, journal mode, and integrity. Check whether retention actually returns disk space.
  • Take consistent backups. For SQLite, use a database-aware backup method or stop writers before copying. Verify the backup instead of assuming a file copy is usable.
  • Create real quiescence. Pause or stagger scheduled work, stop workers that hold maintenance targets open, and avoid parallel status probes during migrations.
  • Validate the new schema. Compare the effective configuration with the version being installed. Do not assume an accepted legacy key will be silently translated.
  • Upgrade the stack coherently. Core and official plugins should converge in the same maintenance window, with migrations allowed to finish before repeated restarts.
  • Verify capabilities, not just processes. Check the gateway, messaging channels, Codex execution, and any other critical tools. Then run actual semantic-memory and session-recall queries.
  • Plan rollback and downtime. Know which files and packages would be restored, and document which steps require the gateway to remain offline.
  • Monitor recurrence. Retention, freelist growth, compaction, and index health need ongoing checks. A one-time vacuum fixes a file; it does not fix the policy that allowed the file to grow.

The Maintenance Window Is Part of the Product

Self-hosted AI is often discussed in terms of privacy, control, model choice, and extensibility. Those advantages are real. So is the maintenance burden that comes with persistent local state.

The more useful an agent becomes, the more history and integration it accumulates. That history is not incidental. It is what makes the agent feel continuous. But it also means upgrades must respect data lifecycles, migration boundaries, active writers, and the difference between a service being up and a capability actually working.

Our field report ends well: the database was compacted with integrity intact, the system reached the new version, and memory plus session retrieval were restored and verified. The important result is not that we found one magic fix. It is that the system became understandable again after we separated the incident into layers and tested each layer on its own terms.

A package update is an event. An operational upgrade is a verified state. For a self-hosted AI agent, the maintenance window is what gets you from one to the other.