Documentation
Git instead of a database
"Everything is Git" describes the architectural rule behind Everlock.
This article is the more mechanical version of that idea.
What does it actually mean to build application storage on top of bare Git repositories rather than a database?
The answer is not that Everlock turns SQL into Git commands.
It uses a much smaller storage primitive and accepts the consequences.
That primitive is a path-based store where every durable change becomes a commit.
One storage interface
At the application boundary, the storage API looks roughly like this:
read(path) -> bytes or None
write(path, bytes) -> CommitId
delete(path) -> CommitId
commit(ops) -> CommitId
list(prefix) -> [paths]
history(path) -> [HistoryEntry]
read_at(path, id) -> bytes at a past commit
Each module opens a store by name.
Each store is an ordinary bare Git repository on disk:
data/<name>.git/
Not a custom format that happens to use hashes.
Not a database with a Git export layer.
A normal bare repository.
That is important because it means the representation remains understandable outside Everlock.
Paths become keys
The store is deliberately simple.
A path identifies a value:
contacts/alice.vcf
calendar/event-123.ics
config/frontend.toml
The value is bytes.
The current tree gives the current set of values.
That means the application can reason in terms of files and paths while Git stores the immutable objects underneath.
The model is not good for every workload.
It is, however, a surprisingly good fit for a lot of human-scale self-hosted data.
Every durable write has history
A normal key-value store might say:
write(path, bytes)
and the old value simply stops being current.
Everlock's write returns a commit id:
write(path, bytes) -> CommitId
That return value matters.
The write did not merely mutate storage.
It produced a new named state in the repository history.
So history is not an optional audit feature that each backend has to remember to update.
There is no durable write path that bypasses history.
That removes an entire category of application-level bookkeeping.
Batch changes become one commit
Some application operations touch more than one path.
For example:
write record
update index
change metadata
If each operation became a separate externally visible state, readers could observe an intermediate version that does not make sense.
The store therefore supports batching:
commit(ops) -> CommitId
Several file changes become one Git commit.
Conceptually:
state A
|
| change path 1
| change path 2
| delete path 3
v
state B
The commit is the application-level persistence boundary.
It is not a general replacement for serializable database transactions, but it gives coherent multi-file state transitions inside one store.
History becomes a storage operation
Because the repository already contains previous versions, the store can expose operations such as:
history(path)
read_at(path, commit)
That makes "what did this file look like last Tuesday?" a storage question rather than a feature each application module has to implement.
The same is true for auditing.
The commit graph already records changes.
The storage layer does not need a separate changelog table that can accidentally get out of sync with the data it describes.
This is one of the main reasons the design is attractive.
Versioning is not beside the data.
It is how the data is stored.
The obvious missing feature is querying
A relational database gives you an enormous amount of machinery:
SELECT ...
WHERE ...
JOIN ...
ORDER BY ...
The Git store gives you paths and bytes.
That is a very different tool.
If Everlock needs to find every record matching some property, Git does not magically provide an index for that query.
The application has to read records or maintain a derived index.
That cost is real.
The important architectural rule is that the index should remain rebuildable from the repositories.
Then the repository stays authoritative while the index remains an optimization.
Git repository -> durable truth
index -> rebuildable view
Once the index contains information that cannot be reconstructed from Git, the architecture has quietly acquired a second durable store.
Concurrency needs an explicit answer
The other obvious challenge is concurrent writes.
Consider:
read current value
make decision
write new value
Two writers can both read the same state and then overwrite one another.
A database usually solves this with transactions or optimistic concurrency mechanisms.
Everlock's store uses a guarded write.
The caller can provide a precondition based on the version it observed:
let seen = Precondition ;
store.commit_checked?;
The precondition is checked while holding the same lock used for the commit.
There is no gap between:
is the state still what I read?
and:
commit my change
The second writer gets a PreconditionFailed instead of silently replacing a newer value.
This is optimistic concurrency control
The model is familiar even if the implementation is not SQL.
A writer says:
Apply this change only if the state still matches what I observed.
That is enough for many application workflows.
It is not the same as serializable transactions across arbitrary stores.
Everlock does not pretend otherwise.
Inside one store, changes can be committed atomically.
Across two independent repositories, there is no single atomic transaction joining them.
That is one of the boundaries the architecture has to respect.
Every write being a commit changes workload design
Git commits are cheap relative to many heavyweight persistence operations.
They are not free.
A workload that produces enormous numbers of tiny independent writes will feel that overhead.
So batching is not merely a micro-optimization.
It is part of the intended write model.
If an operation naturally contains several related changes, they should become one commit.
For very high-frequency event data, the right boundary may be even larger.
This is exactly why access logging becomes interesting: an HTTP request is often too small a unit to deserve its own commit.
The store design works best when application-level persistence has meaningful boundaries.
Growth is a consequence, not an accident
The history property also means stores grow.
If every old state remains part of repository history, then space usage accumulates.
That is expected behavior.
Eventually retention has to answer how much history the system should preserve.
In Everlock, that problem is particularly delicate because deleting history pushes directly against the reason Git was chosen.
The storage model therefore makes retention an architectural decision rather than a background housekeeping detail.
That awkwardness is the cost of history being real.
Backup becomes simpler because representation is ordinary
The store's on-disk representation has another important consequence.
A backup system does not need to ask Everlock to translate the data into an export format.
The store is already a Git repository.
That means normal Git and filesystem tooling can copy it.
The details of a robust backup still matter.
A mirror is not automatically a point-in-time backup, and a tested restore matters more than a successful copy job.
But the backup unit itself is straightforward.
There is no second serialized form of the application state.
Replication uses the same object model
The same applies to moving data between machines.
A repository already consists of content-addressed objects and refs.
Git already knows how to negotiate which objects another repository needs.
So Everlock does not need to define another wire format merely to reproduce store state elsewhere.
Again, this does not solve conflicts, authorization, or application semantics.
It removes one layer of transport design.
That is often enough to simplify the system substantially.
Integrity comes from the object model
Git names objects by hashes derived from their contents.
That gives the repository built-in integrity properties.
If an object does not match its expected id, something is wrong.
Everlock does not need to invent another per-record checksum format to get basic content-addressed verification.
This does not make repositories immune to every form of corruption.
It means integrity is already part of the storage format rather than an optional application feature.
Why this suits Everlock's data
The trade makes the most sense for data with a particular shape.
Personal photos can be large in bytes without being high-frequency transactional data.
Calendars contain relatively few records and benefit strongly from history.
Contacts are small.
Configuration changes rarely and is exactly the kind of thing where previous states are valuable.
Mail is largely append-oriented.
These workloads are not zero-cost in Git.
They are simply close enough to Git's strengths that the benefits of a single versioned representation are valuable.
A high-write-rate analytics system would be a very different story.
What the database would have bought
Using PostgreSQL would immediately provide better querying, indexes, constraints, transactions, and mature operational tooling.
Those are significant advantages.
But Everlock would then have two durable representations to manage in many backends:
database records
files or binary objects
The relationship between them becomes application logic.
Backups have to preserve a consistent pair.
Migrations have to move both.
History has to be added deliberately.
An export path becomes necessary if the user wants a representation independent of the application.
The Git store gives up some database capability in exchange for collapsing those concerns into one durable format.
That is the actual trade.
The store is small on purpose
The useful abstraction is not:
Git pretending to be PostgreSQL
It is:
paths + bytes + commits
From that, Everlock gets a coherent set of properties:
current state
history
time travel
batch commits
content identity
replication
ordinary backups
external inspectability
What it does not get must either be built above the store or accepted as a boundary.
That makes the architecture easier to reason about.
The database is not wrong, it is unnecessary here
"Git instead of a database" sounds more ideological than the actual design is.
There is no claim that Git is a universally better database.
For many applications it would be an awful one.
The point is that Everlock's persistent data is mostly human-scale, document-like, and valuable to keep as history.
For that shape, Git provides several properties the project cares about at once.
And most importantly, the representation outlives Everlock itself.
If the software disappears, the stores are still bare Git repositories full of ordinary files and ordinary history.
No migration is needed to make them portable.
They were portable from the beginning.
See versioned storage for the reference, and architecture for how the backends sit on top of it.