lucasoshiro3 minutes ago
In newer versions of Git you check this by using `git repo structure`
time4tea7 minutes ago
du -sh
Human readable, no explanation required.
cocoto8 hours ago
> The du options are: s to summarize and b to display bytes.
Small nitpick: Use long options and your code samples become self explanatory!
zetanor5 hours ago
On the flip side, not everyone uses GNU coreutils (though I do on every machine that runs Git), and I wouldn't have been confident that "du -c" and "du --total" are fully interchangeable prior to reading the manual just now. For Git, he could have confidently used "git --message" instead of "git -m", but then again, I'm sure it would have surprised many Git users ("is that -m?").
analog_daddy7 hours ago
This is one of my biggest pet peeves! I hate reading shell scripts which uses small options. It is okay to use short options when interactively using shell, but there is no excuse to not spend time and try converting a small options shell script to use long options if its going to be read by someone else, especially if you anyway have to write comments explaining it. The biggest part i hate is that, some utilities don’t have long option counterpart!
NetOpWibby8 hours ago
I always use long options when possible. I don’t need terseness, computers do. I wanna help future me figure out how to run something.
nathanpankon9 hours ago
I played with making a git alternative that was streamlined according to workflow. I haven't completely abandoned it, but my core idea was that you can use ast path to compress the data further.
Like you said git is really efficient, and even though I went in with some criticism because of the work(mess) that is chromium and the deep hacking I did there to fix stuff between versions without forcing a complete rebuild
I came to the conclusion that as a compression alternative, my approach wasn't worth it. Git did a better job!
Anyway here's the attempt: https://github.com/pankon/gat
Anduia5 hours ago
Very interesting! Maybe the AST part is more useful for diffs and merges than for compression.
nathanpankon3 hours ago
There are some attempts at this (semanticdiff, etc) and they seem very good. But I was looking for a tool that could actively edit a set portion of code, maintaining a small diff above a moving target
In that case merging from "main" starts to be extremely difficult
Anduia9 hours ago
It is not public is it?
nathanpankon8 hours ago
Good call, made it public :)
Thanks!
alansaber2 hours ago
Great name btw lol
jamesblonde6 hours ago
Git was designed with assumptions about local file storage which make it challenge when storing data in a distributed file system. We built our distributed filesystem, HopsFS, to store data in S3. To support high performance git in FUSE, our writes hit remote NVMe disk(s) and asynchronously sync to S3 (you can replicate on NVMes or just do failure recovery). Previously, we used remote NVMe disks as a write-through cache, but that doesn't work with git, S3 latency is too high.
EngineeringStuf4 hours ago
We built similar at Swarmfile, and it's compatible with Git too
CodesInChaos7 hours ago
Does git support better compression algorithms than deflate, like zstd or brotli?
masklinn6 hours ago
No. You can configure the compression level but that's at far as it goes.
You can configure loose object and pack compression separately (core.compression and pack.compression) so loose object compression could be switched to an other algorithm as it's completely local[1] (likely a low-compression zstd), but pack files are part of the network exchange, so things are a bit more complicated there. Also an annoyance there is that git uses raw zlib there is no identifying header or anything.
There are old threads on the mailing list, as well as a gitlab issue.
Although you can also tune compression per pack when you create them explicitly, on $DAYJOB's repository in order to speed up gc some I have .keep packs which segregate images uncompressed in their own packs, because there's little point wasting cycles on zlib-compressing jpegs and pngs.
[1]: unless you're using the "dumb http" protocol, or performing (n)fs-mediated clones
CodesInChaos5 hours ago
I'm a bit surprised Github or some other big git hoster didn't push for such a protocol improvement. Since I assume that smaller and cheaper to create packs would reduce their costs noticeably.
masklinn3 hours ago
Packed objects are compressed individually. I expect the costs and benefits of compression are generally dwarfed by delta-ification.
fiddlerwoaroof10 hours ago
It can get even smaller when loose objects get packed into packfiles because of delta compression: a whole bunch of identical files can be stored in minimal overhead over storing one copy.
gblargg7 hours ago
Right, packfile size is really what matters in the long-term. A tiny change to a file in the short term just adds the whole thing again (compressed), whereas a tiny change in a packfile shouldn't add much at all. Unless you've got very limited disk space, the long-term is the only size that matters.
masklinn6 hours ago
> whereas a tiny change in a packfile shouldn't add much at all
Well to the extent that git is able to find a suitable delta, which is the difficult part.
bewuethr8 hours ago
I think the 250 files all containing just "foo" are stored as a single blob, too, even before creating a packfile.
fiddlerwoaroof7 hours ago
I forgot about that, delta compression still can save quite a bit when you have nearly-identical files.
bewuethr7 hours ago
Exactly, a tiny change in a large file produces a full new blob, but is stored very efficiently as a delta in a packfile.
grumbelbart29 hours ago
Did you do —apparent-size?
SushiHippie6 hours ago
The -b option is equivalent to '--apparent-size --block-size=1'
dhruv_kernlai4 hours ago
[flagged]
llj12310 hours ago
[dead]