Documentation
Everything Is Git
Everlock starts with a fairly unreasonable constraint:
Everything that matters should live in Git.
Not just the source code for Everlock. The data managed by Everlock.
Photos. Calendars. Configuration. Application state. Permissions. Content served by the system. If Everlock is responsible for something durable, the default question is whether that thing can be represented inside a Git repository.
This is not because Git is a particularly fashionable database. It isn't a database at all in the usual sense.
The reason is simpler: I wanted the durable state of a self-hosted system to remain understandable and useful without the application that created it.
That decision has consequences. Some are extremely convenient. Some are awkward enough to deserve their own articles. But most of Everlock's architecture starts here.
The repository is the product state
In a typical application, Git disappears after deployment.
git repository -> build -> application
Once the application is running, its state usually goes somewhere else:
application
|
+--> PostgreSQL
+--> object storage
+--> local filesystem
+--> Redis
Git describes how the application was built. It does not describe what the application currently contains.
Everlock deliberately collapses that distinction. The running software operates on repositories, and those repositories are the durable state.
If Everlock disappears tomorrow, the repositories remain.
That property is more important to me than making every individual storage operation as easy as possible.
Why Git?
Git is a strange choice if the question is:
What is the best database for this application?
But that isn't the question Everlock starts with.
The question is:
What properties should the user's data have after Everlock has written it?
I want the answer to include:
copyable
versioned
inspectable
portable
replicable
recoverable
independent of Everlock
Git already provides an unusually useful combination of those properties.
A file is not just a file
Suppose Everlock stores a calendar entry as:
calendar/event.ics
With a normal filesystem, the current file is the state. If it changes, the previous state disappears unless another system preserves it.
Inside Git, the same file participates in history:
A -> B -> C
At A, the event had one value. At B, its time changed. At C, its
description changed.
The application doesn't need to invent a separate version table just to preserve those states. Versioning is part of the storage model.
A commit can also say that several file changes belong together. That is much stronger than merely having timestamps on individual files.
The commit is an application primitive
Once application state lives in Git, a commit stops being something only developers create. It becomes an application-level transaction boundary.
Imagine an operation changes:
calendar/event.ics
calendar/index.json
metadata/last-modified
Everlock can prepare those changes and create one commit:
state A -> commit -> state B
That does not make Git a replacement for every database transaction. But for many document-oriented changes, the model is useful.
The commit gives the state a durable identity, and other Git machinery immediately understands that identity.
History comes without a second history system
Applications frequently end up implementing history themselves. A table begins with a current value and timestamps, then somebody asks for previous versions, so revisions or history tables appear.
Git already has a graph of immutable objects designed around versions of trees.
Everlock can use that graph rather than building another one.
This leads to a recurring architectural rule:
If Git already has a primitive with the right semantics, use the Git primitive.
Not because every problem is secretly source control, but because fewer parallel representations of the same state means fewer things that can disagree.
The data can leave without asking Everlock
This is probably the most important property.
A self-hosted application can still mean:
your server
your disk
our private database format
our application required to understand it
The physical location changed. The dependency did not.
With Everlock, the repository itself should remain useful.
It can be cloned, mirrored, inspected with Git tooling, and individual files can be extracted without asking Everlock to export them first. History can be traversed without a running Everlock server.
If the software stops working, the first recovery step does not have to be:
Get Everlock working again so I can get my data out.
The repository is already there.
Replication is a protocol we already have
Once durable state is a Git repository, moving that state between machines becomes a Git problem.
Git is very good at answering:
Which objects do you have, and which objects do I need to send you?
Instead of inventing an Everlock synchronization format:
Everlock A Everlock B
objects + refs <--- Git ---> objects + refs
The application data might be photos rather than source code. Git doesn't care.
This doesn't mean every synchronization problem disappears. Conflicts, authorization, large objects, retention, and application semantics still exist.
But the transport does not need to be invented from scratch.
Backup becomes pleasantly boring
The same property applies to backups.
If repositories are the durable state, a backup system does not need to understand the Everlock data model. It can copy repositories.
Everlock understands the data.
Git understands the repository.
The backup system copies the repository.
There are still important distinctions between replication and real backup. A synchronized mistake is still a synchronized mistake, which is why point-in-time snapshots and restore testing matter.
But the backup does not need an Everlock-specific export format.
The unit being protected is already portable.
Git gives us names for alternative states
Branches and refs become interesting once the repository contains application data.
A ref is simply a name pointing into the object graph:
refs/heads/main
refs/pulls/123/head
One can mean authoritative state while another means proposed state.
That makes features such as pull requests surprisingly natural. A proposed change does not need to be copied into a separate database representation.
It can remain:
Git objects + a ref
Everlock adds policy and meaning around the primitive.
Then you encounter data that does not fit
This architecture becomes most interesting when Git is a bad fit.
Access logs are a good example.
Creating a commit for every HTTP request would be technically possible and architecturally ridiculous.
An access request is too small and too frequent to deserve its own Git commit. But abandoning Git entirely would introduce another durable storage system.
So event granularity and persistence granularity can be separated:
request
request
request
request
|
v
in-memory batch
|
v
log segment
|
v
Git commit
Git remains the durable representation. It simply does not participate in every individual event.
That distinction matters:
Everything durable can live in Git without every operation becoming a Git operation.
Git is not the hot path by definition
Using Git as durable storage does not require forcing Git synchronously into every request path.
Everlock can process something in memory, build an intermediate representation, batch changes, and commit when a meaningful persistence boundary has been reached.
The architectural constraint is about where durable truth ends up.
It is not a commandment that every function call must immediately create an object and update a ref.
Git does not make every workload efficient
There are obvious tradeoffs.
Git is not a high-frequency event database. It is not a relational query engine. It does not provide arbitrary secondary indexes. Large binary files behave differently from small text files. Garbage collection and repository maintenance matter.
History can also become a disadvantage when data must actually disappear:
deleted from current tree
is not the same as:
removed from repository history
Those are not reasons to pretend Git is something it isn't.
Everlock does not use Git because Git is optimal for every storage workload. It uses Git because the properties of the resulting durable state are valuable enough to accept some awkwardness around particular workloads.
The constraint is useful because it is inconvenient
Architecture becomes easier if every new requirement can introduce the perfect specialized service.
Need search? Add a search database.
Need events? Add an event store.
Need files? Add object storage.
Need metadata? Add PostgreSQL.
Need cache? Add Redis.
Each decision can be reasonable on its own. The resulting system can still end up with five definitions of durable state.
Then backup means coordinating all five. Migration means moving all five. Recovery means restoring all five to compatible points in time.
A constraint prevents some of that expansion.
When a new Everlock feature needs persistence, the first question is not:
Which database would be nicest for this?
It is:
How far can this fit into the repository model without becoming unreasonable?
Sometimes that produces a clean design. Sometimes it produces a compromise such as batching. Eventually there may be cases where Git is the wrong place.
But that decision should be earned rather than assumed.
The repository becomes the boundary
This is the mental model I find most useful:
protocols
HTTP / Git / mail / ...
|
v
Everlock
|
v
application semantics
|
v
Git repositories
|
v
durable truth
Protocols can change. The UI can change. The Everlock implementation can change. Indexes can be rebuilt. Caches can disappear. Temporary queues can be lost.
The repositories are the boundary after which the data is supposed to remain.
That gives us a useful test for new components.
Does a new database contain state that cannot be reconstructed from the repositories? Then it is no longer merely an index.
Does a queue contain information that would be permanently lost if the process crashes? Then we need to decide whether that information belongs in durable state.
Does a feature require an export operation before its data becomes independently useful? Then perhaps the representation is too application-specific.
Everything is Git, but Git is not everything
"Everything is Git" is deliberately provocative.
Taken literally, it would be terrible architecture.
CPU caches are not Git. Network connections are not Git. Search indexes do not have to be Git. A request currently being processed does not have to be Git.
The useful version is narrower:
Everything that Everlock considers durable, authoritative user state should have a representation in Git.
That leaves plenty of room for ephemeral systems around it:
cache rebuildable
index rebuildable
queue temporary
runtime state temporary
Git repositories authoritative
Once that distinction is clear, Git stops looking like an attempt to use source control as a universal database.
It becomes the durable boundary of the application.
Start with the thing you want to keep
Most application architectures begin with the running system.
Which database should the server use? Which API should it expose? Which services should be deployed?
Everlock starts one layer lower.
If the server disappears, what should still be there?
My answer is a collection of ordinary Git repositories containing the state and its history.
Everything above that can be rebuilt.
That single constraint is responsible for many of Everlock's convenient properties and several of its awkward ones.
It is why backups can operate on repositories.
It is why proposed changes can be refs.
It is why repository permissions matter.
It is why high-frequency access logs need batching instead of one commit per request.
And it is why the application does not need to be the only program capable of understanding where the user's data lives.
So when Everlock says everything is Git, it is not really making a claim about Git.
It is making a claim about ownership:
The durable state should outlive the software that manages it.