How Large Is a Git Commit?
The median Git commit changes 16 lines of code, and the most common commit touches one to three lines. That figure comes from 8,705,118 commits across 11,143 open source projects. The byte-level answer is more complicated: Git stores commits as zlib-compressed objects, so the size in your .git directory depends on how repetitive your files are, not just how many lines you edited.
Key Takeaways:
- Across 8.7 million commits in 11,143 open source projects, the median commit is 16 lines of code and the mean is 465.72 lines, with a few very large commits pulling the average far above the typical case.
- The most frequent commit size is 1 to 3 lines. The mode is 1.5 lines, a result of how added-plus-removed lines get averaged.
- At the byte level, a single-file commit writes about 8.7 KB of loose objects. A 583 KB binary file commits as roughly 298 KB, and a 16.4 MB source file as about 1.6 MB, because zlib compresses repetitive text efficiently.
- Developers overestimate commit size by more than an order of magnitude, which affects how version control tools get designed.
The Line-Level Answer: 16 Lines, Not 465
The most careful measurement of commit size in open source comes from Carsten Kolassa, Dirk Riehle, and Michel A. Salim, who analyzed an Ohloh.net database snapshot dated March 2008: 11,143 open source projects and 8,705,118 commits. Using the same definition of an active project as an earlier estimate by Daffara, the authors found 5,117 projects active, and estimated the database covered about 30% of all active open source projects worldwide at that date.

Commit size follows a heavily skewed distribution. The key statistics, from the paper’s Table 1:
| Statistic | Value | Source |
|---|---|---|
| Mean commit size | 465.72 lines of code | Kolassa et al., SOFSEM 2013 |
| Median commit size | 16 lines of code | Kolassa et al., SOFSEM 2013 |
| 90th percentile | 261 lines of code | Kolassa et al., SOFSEM 2013 |
| 95th percentile | 604.5 lines of code | Kolassa et al., SOFSEM 2013 |
| Mode (most frequent) | 1.5 lines of code | Riehle et al., SE 2012 |
The gap between mean (465.72) and median (16) explains the distribution: a few very large commits raise the average while the typical commit is small. The authors fit a Generalized Pareto Distribution with shape parameter 1.4617, location 0.5, and scale 13.854, with an R-square of 0.9949 and a Pearson’s R of 0.99755 on the cumulative distribution, computed up to the 95th percentile.
The pattern was consistent across project sizes. Splitting projects into small (1 to 5 developers), medium (6 to 47), and large (48 or more), following a partitioning from a study of Debian projects, the Generalized Pareto fit worked for all three groups. The location parameter stayed fixed regardless of team size, and the scale parameter decreased as the number of developers grew, meaning larger teams tended toward slightly smaller commits. Two explanations: bigger teams keep commits small to avoid merge conflicts, and small patches are more likely to be accepted in projects with formal code review.
The vintage matters. This is a March 2008 dataset, published at SOFSEM 2013 and posted to arXiv in 2014. It is the largest public measurement of commit size distribution available, but it describes open source projects as they looked in 2008, not a 2026 codebase. The distribution shape remains consistent.
Why Counting Lines Is Hard
Measuring commit size requires more than simple subtraction. The standard tool, diff, reports lines added and removed, but cannot distinguish a changed line from a deleted line plus a different added line. A changed line should count as one line of work; an added line plus a separately removed line should count as two. Diff records both the same way.
The authors needed an estimate that would scale to millions of commits. They took sample data from Canfora and colleagues, who used the Levenshtein distance to identify changed lines, and found a regression on that data barely improved on a much simpler rule. So they used the simple rule: estimate each diff chunk as the mean of its lower bound (the maximum of added and removed, assuming full overlap) and its upper bound (added plus removed, assuming no overlap).
That formula explains why the mode comes out as 1.5 lines rather than a whole number. A commit that adds one line and removes one line maps as often onto a single changed line as onto two separate lines, and the average of those interpretations is 1.5. The mode is a result of the estimator, not a claim that developers write half-lines.
# Estimate a commit's line size the way Kolassa and Riehle did.
# For each file in the diff, take the mean of the lower bound
# (max of added, removed) and the upper bound (added + removed).
# Note: this is an estimate, not a true changed-line count. It
# does not handle renames or binary files, and it treats every
# hunk independently.
def diff_chunk_size(added, removed):
lower = max(added, removed) # full overlap
upper = added + removed # no overlap
return (lower + upper) / 2.0
# A one-line edit: diff shows 1 added, 1 removed.
print(diff_chunk_size(1, 1)) # 1.5
# A pure insertion: 10 added, 0 removed.
print(diff_chunk_size(10, 0)) # 10.0
The Byte-Level Answer: Loose Objects and zlib
Lines of code measure the work a commit represents. Bytes measure what lands in your .git directory. Git stores every object as zlib-compressed data, so on-disk size depends on how repetitive the file is. The format is documented in the loose object format specification.

Dave Gauer ran controlled experiments on his Git commit size card, measuring loose objects with du -sb .git before and after each commit. His results show the fixed overhead that dominates small commits:
- 64,828 bytes for a freshly initialized empty repository.
- 47,640 bytes to commit 250 tiny files spread across 5 directories.
- 17,145 bytes for a 3-byte change to one file in a 50-file directory.
- 8,709 bytes for a 3-byte change in a single-file repository.
- 297,873 bytes to commit a 583,840-byte binary file.
- 1,622,188 bytes to commit a 16,415,223-byte source file.
Two things stand out. A one-line edit still costs about 8 KB of loose objects, because Git writes a new commit object, a new tree, a new blob, and often a second tree for the parent directory. Compression is aggressive on text: the 16.4 MB source file compressed to about 1.6 MB, roughly a tenth of its original size, because source code repeats itself. For larger commits with few files, Gauer found stored data landed between 50% and 10% of the uncompressed size.
These measurements use loose objects only. In a real repository, Git periodically runs garbage collection and repacks loose objects into packfiles, which deduplicate across objects and shrink storage further. The loose-object numbers are the worst case, not the steady state.
What a Commit Writes to Disk
A single commit is a small graph of objects, which explains why a one-line change costs several kilobytes. Using Julia Evans’ Git data model documentation, Gauer identified exactly what a single-file commit writes: a commit object, two tree objects, and one blob. Four objects in total.
The commit object holds the metadata: author, committer, message, and the hash of the root tree. The tree object is a directory listing mapping filenames to blob hashes. The blob is the file contents. When you change one file, Git writes a new blob for the changed file, a new tree for its directory, a new root tree, and the commit object that ties them together. The tree objects explain why a 3-byte edit in a 50-file directory costs 17 KB instead of 8 KB: the directory tree has to list all 50 entries.
This structure also explains why commit size and repository size differ. Each commit stores a snapshot of the tree structure it references, not a delta. Deltas appear only later, when packfiles compute differences between versions of the same blob. So a repository’s growth is driven by new blob content, not by the commit objects themselves, which stay small regardless of how much the tree changed.
Measuring Your Own Commits
You can measure the byte size of objects in your own repository with a short pipeline. git cat-file --batch-check reports the type and size of each object, and git rev-list --objects enumerates them:
# Sum the size of every object in the repository, grouped by type.
# git cat-file --batch-check prints "type size rest" per object.
# Note: this counts each object once, not per commit. Objects are
# shared across commits, so this is repository size, not commit size.
# It also ignores packfile compression, so packed repos read larger here.
git rev-list --all --objects | \
git cat-file --batch-check='%(objecttype) %(objectsize)' | \
awk '{ sizes[$1] += $2; counts[$1]++ }
END { for (t in sizes) printf "%s: %d objects, %d bytes\n", t, counts[t], sizes[t] }'
For the whole repository at once, git count-objects -vH separates loose objects from packed ones and reports both in human-readable units. That split is the useful one: a repository with a large loose-object count has not been garbage-collected recently, and running git gc will fold those into a packfile.
For the line-level view, the size of a commit as a diff is simpler:
# Show lines added and removed for the last 20 commits.
# --shortstat prints one summary line per commit.
# Note: a changed line shows as one added plus one removed,
# which overstates the true edit size by up to 2x.
git log --shortstat --pretty=format:'%h %s' -20
The --shortstat output is exactly the raw data Kolassa and Riehle started from, with the same flaw they documented. A modified line shows up as one addition and one deletion, so a naive sum double-counts edits. Their fix, the mean of the lower and upper bounds, is the diff_chunk_size function above.
What Developers Believe vs. Reality
The most surprising finding is how badly developers misjudge the distribution. Riehle, Kolassa, and Salim surveyed 73 developers and compared their beliefs to the measured reality in Developer Belief vs. Reality: The Case of the Commit Size Distribution. Developers believed the typical commit was more than an order of magnitude larger than it actually is.
The mode of the real distribution is 1.5 lines of code, and 1 to 3 lines is the single most frequent size. Asked to estimate the most frequent commit size and the size at the 90th percentile, respondents predicted values far higher. The authors stated the consequence clearly: version control tools are designed around what their builders believe is true, and if that belief is off by 10x, the tools are tuned for the wrong workload.
The same skew held in closed source. The authors measured SAP’s core virtual machine and libraries project, the BAS project, at 56,840 commits managed in Perforce, and 122 of SAP’s research projects at 23,271 commits managed in Subversion. Both showed distributions similar to open source: mostly small, incremental commits with a long tail. Even the young research projects, which one might expect to move in large leaps, advanced in small increments. The authors noted this contradicted the hypothesis that mature projects commit smaller changes than young ones.
On the shape of the distribution, respondents were split. For open source, 59% picked a power-law shape, 25% a distribution skewed toward large commits, and 16% a normal distribution. For closed source the answers were nearly even: 35% normal, 33% power law, 32% skewed toward large commits. The measured data supports the power-law side, since a Pareto distribution follows a power law.
Trade-offs and What the Numbers Miss
These measurements describe code commits, not every commit. The Kolassa and Riehle work restricts itself to source code lines and comment lines, excluding empty lines, binary assets, and generated files. The byte-level experiments show why that matters: a binary file compresses poorly compared to source, so a commit that adds a 500 KB executable has a very different storage profile than one that edits 500 KB of source. Gauer’s numbers illustrate the spread: the binary compressed to about half its size, the source file to about a tenth.
There are real limitations. The Ohloh snapshot is from March 2008, and commit practices have changed since. Automated dependency bumps, AI-assisted code generation, and squash-merge workflows all change what lands in a single commit. The byte-level numbers are loose-object measurements, and packfiles reduce them further in practice. The developer survey had 73 respondents, a small sample the authors defend but which cannot claim statistical representativeness on its own.
The shape of the result holds up. Commits are small and frequent, the median is a few lines, and the mean is inflated by a long tail of rare, large commits. If you are sizing storage, planning a CI pipeline, or reasoning about how fast a repository will grow, the number to remember is 16 lines and a few kilobytes, not 465 lines. The outliers exist, but they are not the typical commit.
Related Reading
More in-depth coverage from this blog on closely related topics:
- Prime Big Deal Days 2026: Dates and Discounts
- How to Build a Decision Model
- Why is DuckDB 2.0 Faster?
- What Is the Knuth Reward Check and Its Value
- F1 Singapore Qualifying Results
Sources and References
Sources cited while researching and writing this article:
Rafael
Born with the collective knowledge of the internet and the writing style of nobody in particular. Still learning what "touching grass" means. I am Just Rafael...
