Documentation

Last updated: 2026-09-12

If Everything Is in Git, Where Do the Access Logs Go?

Everlock stores its state in Git repositories.

That is a useful constraint.

Files are versioned. Changes have history. Replication can use Git. Backups can use Git. A repository can be inspected without needing a separate database or a special Everlock export format.

Then I wanted to add an access log.

At first this sounds almost embarrassingly simple. A request arrives, so append a line to a log.

But if Git is the persistence layer, "append a line" is no longer a small operation.

It raises a surprisingly useful question:

Does every piece of persistent data deserve to be a commit?

For an access log, the answer is almost certainly no.

But saying no creates another problem. If the log still has to live in Git, what exactly should be stored?

A request is much smaller than a commit

Consider a normal HTTP access log entry:

2026-09-02T10:42:17Z GET /photos/IMG_1234.jpg 200 184291

The interesting information might only be a few dozen bytes.

Representing that request as its own Git commit would require considerably more structure.

Conceptually, each request would become something like:

request
  |
  v
new log content
  |
  v
blob
  |
  v
tree
  |
  v
commit
  |
  v
ref update

Git is very good at storing lots of objects, but that doesn't mean creating a commit for every HTTP request is a sensible use of it.

A busy service could create thousands or millions of commits whose only purpose is to record individual events.

The history would be technically correct and practically absurd.

The problem is granularity

This isn't really a Git performance problem.

It is a mismatch between two different kinds of data.

Most of Everlock's normal state changes at human or application scale:

photo added
calendar changed
configuration updated
repository permission changed

An access log changes at event scale:

request
request
request
request
request
request

Git naturally gives meaning to a change boundary.

A commit says:

This is a state worth naming.

That is useful for a configuration change.

It is much less useful for the fact that somebody requested /favicon.ico.

The mistake would be assuming that because Git is the persistence mechanism, every incoming event has to map directly to Git's transaction boundary.

It doesn't.

Persistence and commits are not the same thing

This distinction took me a while to formulate clearly.

When I say that Everlock stores its data in Git, there are actually two separate properties hiding inside that statement:

  1. The durable representation belongs in a Git repository.
  2. Every mutation must immediately create a Git commit.

The first is important to Everlock.

The second is not.

That opens up a much more useful design space.

An access request can first exist as an event:

HTTP request
     |
     v
compact log event
     |
     v
memory queue

Several events can then be collected:

event
event
event
event
event
  |
  v
batch

And the batch can become one persistent Git change:

batch
  |
  v
log object
  |
  v
commit

Git remains the durable store.

It simply stops being part of the request path.

Keep Git out of the hot path

This matters for more than object count.

An HTTP request should not have to wait for a Git transaction just because the server wants to record that the request happened.

The access path and persistence path have very different latency requirements.

The request path looks roughly like this:

request
   |
   +--> handle request
   |
   +--> emit log event

The logging path can be independent:

log event
   |
   v
queue
   |
   v
batch
   |
   v
persist to Git

In Rust, an in-memory channel fits this model naturally. Request handlers can send compact events into an mpsc queue while a separate task consumes them and periodically persists a batch.

That gives the request handler a cheap operation and gives the persistence layer something large enough to justify a commit.

It also introduces an uncomfortable property:

The events in the queue are not yet durable.

How much logging can you afford to lose?

Once events are buffered in memory, a crash can lose the current batch.

That sounds bad, but the important question is not whether data can be lost.

The question is what durability the data actually requires.

For some audit logs, losing even one event is unacceptable.

For a generic HTTP access log, losing the last few seconds during a process crash may be entirely reasonable.

Those are different products with different guarantees.

This is another place where using Git can force the requirement to become explicit.

If I want every event durably recorded before responding to the request, then an in-memory queue isn't sufficient. I need some form of synchronous durable journal.

But now I have introduced another persistence system solely to protect data before eventually putting it into Git.

That might be correct for an audit system.

For an access log, it may be solving a problem I don't actually have.

Batching needs a boundary

Once requests are collected, the next question is when to flush them.

There are several obvious boundaries:

every N events
every N seconds
when memory reaches N bytes
on graceful shutdown

In practice, a combination is more useful than picking only one.

A quiet server shouldn't keep a handful of log entries in memory for hours just because the batch size hasn't been reached.

A busy server shouldn't wait for a timer while the queue grows without limit.

So the persistence decision can be driven by both time and size:

flush if batch >= size limit
        OR
flush if oldest event >= time limit

The exact numbers are tuning parameters.

The important architectural point is that these are logging decisions, not Git decisions.

Git only sees the resulting batch.

The format matters more when you batch

Once hundreds or thousands of events are accumulated before persistence, representation starts to matter.

A verbose structure such as JSON is pleasant to inspect:

{
  "timestamp": "2026-09-02T10:42:17Z",
  "method": "GET",
  "path": "/photos/IMG_1234.jpg",
  "status": 200,
  "bytes": 184291
}

But field names are repeated for every request.

For a handful of events, nobody cares.

For millions of events, those repeated bytes become a meaningful part of the data.

Traditional access-log formats solve this by relying on a schema outside each individual event:

2026-09-02T10:42:17Z GET /photos/IMG_1234.jpg 200 184291

The meaning of column four doesn't need to be spelled out on every line.

A binary format can go further, but it also makes the repository harder to inspect with ordinary tools.

That creates a familiar Everlock tradeoff.

The most compact representation isn't automatically the best representation.

If Git is useful partly because people can inspect what is stored, then a slightly larger textual log may be worth the space.

Git compression changes the calculation

There is another reason not to optimize the event format too aggressively.

Git doesn't necessarily store repeated textual data as independent full copies forever.

Its object storage and pack files are designed to compress and delta related content efficiently.

That doesn't make storage free, and access logs are not Git's ideal workload.

But it means the size of the source representation and the eventual repository size are not identical.

A format with repetitive structure may compress surprisingly well.

This makes the first optimization target clearer:

Don't create a ridiculous number of commits.

After that, measure the actual repository before inventing an exotic log encoding.

Should the log be one ever-growing file?

The obvious batched implementation is to append events to something like:

logs/access.log

Each flush produces a new version of that file.

Simple.

But an infinitely growing file has its own drawbacks. Every new version conceptually replaces the previous blob, and tools that inspect the file eventually have to deal with a very large object.

A segmented layout gives the data natural boundaries:

logs/
  2026-09-02/
    10.log
    11.log
    12.log

Or perhaps boundaries based on batch identifiers rather than hours.

Now old segments become immutable.

New requests only affect the current segment.

This maps much better to the append-heavy nature of logging while still producing ordinary files in an ordinary Git tree.

It also gives retention policies something concrete to operate on.

Deleting logs older than a certain period becomes a repository change rather than surgery inside one giant file.

Git history is not retention

There is a catch.

Deleting a log file in Git doesn't immediately remove its contents from the repository.

The old blob remains reachable through history.

That is normally one of Git's best properties.

For log retention, it can be exactly the opposite of what you want.

If the requirement says:

Keep access logs for 30 days.

then committing a deletion after 30 days doesn't actually satisfy that requirement if the old data remains in Git history indefinitely.

This is an important distinction:

deleted from current tree
          !=
removed from repository

Real deletion may eventually require pruning history or designing the log storage so expired data can become unreachable and be garbage-collected.

At that point, retention becomes part of repository maintenance rather than simply file management.

Again, Git makes the semantics difficult to ignore.

Not every event deserves history

The useful conclusion for Everlock wasn't that access logs don't belong in Git.

It was that they don't deserve the same history granularity as normal application state.

A request is an event.

A batch of requests is a reasonable persistence unit.

A commit records that persistence unit.

That gives us three different scales:

request       -> event
many events   -> log segment
log segment   -> Git commit

Trying to collapse those into one scale creates unnecessary work.

The same distinction applies beyond access logs.

Metrics, telemetry, synchronization events, delivery receipts, and other high-frequency data often have the same shape.

They may belong in the same durable system as the rest of the application without deserving one transaction per observation.

"Everything is in Git" is a constraint, not a command

I still like the constraint that Everlock's persistent state lives in Git.

It means the repository remains the thing that can be copied, backed up, inspected, and moved.

But constraints become harmful when they are interpreted too literally.

"Everything is in Git" does not have to mean:

Every event immediately becomes a commit.

It can mean:

Git is the durable representation of the system.

Between an HTTP request and that durable representation, there is room for queues, batches, aggregation, and sensible persistence boundaries.

That small distinction turns access logging from a pathological Git workload into a fairly ordinary event-processing problem.

And it preserves the property that mattered in the first place:

When the server is moved or backed up, the logs move with the repository too.

updates git storage logging design