Rendered at 05:34:48 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
kpcyrd 12 hours ago [-]
This article is full of mistakes and misleading claims:
1) It's claiming SHA1 insecurity is theoretical, while SHAttered from 2017 was specifically a pratical proof of concept. The only reason Git wasn't affected, is because they didn't bother bruteforcing a git-blob prefix.
2) It's claiming collision attacks don't matter, only second-preimage attacks do. This is incorrect, collision attacks are enough for code-smuggling problems, when two repositories are on the same git commit (verified by the full commit hash), yet contain different code in their git checkout.
3) The Linus quote "The real security is in distribution" is arguing that "git's content-addressed system should not be used to address content". It's arguing that, in case of curl|sh, you shouldn't use a sha256sum-gate to pin the content to something you've reviewed, you should instead ensure curl is fetching from an https server.
schacon 12 hours ago [-]
1) I link to the SHAttered paper, as well as Shambles. Git projects were not affected because it is an inefficient attack vector. I say it's impractical to exploit, which I think everyone agrees with.
2) I specifically argue that even if both attacks were practical and cheap, it's still not the problem we should be focusing on.
3) Have you read this email (that I linked to)? It is almost the same general message (20 years ago) that this blog post is. It literally goes though a theoretical object replacement attack and how dumb this scenario is and so SHA-1 is fine.
> 1) I link to the SHAttered paper, as well as Shambles. Git projects were not affected because it is an inefficient attack vector. I say it's impractical to exploit, which I think everyone agrees with.
It seems unlikely it will stay that way forever. Typically attacks get more efficient over time as researchers find improvements, not to mention computers getting better.
In 2015 it was estimated to cost $100,000, now the estimate is down to $10,000. Where will it be in 2035?
schacon 11 hours ago [-]
I specifically argue that it doesn't matter if it's $1 and base my argument and solution around that. So it's irrelevant where it is in 2035.
bawolff 9 hours ago [-]
You still mentioned the (1) part, which is what i object to. I agree its not fatal to your argument.
AlfeG 10 hours ago [-]
Even if it costs zero to create. How do You force people to pull from Your repo?
kpcyrd 2 hours ago [-]
The XZ incident would have been so much worse if it would have involved a collision, pushing one object variant to github.com, and one variant to git.tukaani.org.
Then you would have security researchers making conflicting claims depending on which repository they first pulled from, even though they are on the same git commit hash.
axus 9 hours ago [-]
Hack into the system holding a trusted repository, and swap in your variant with same signature?
jurgenburgen 30 minutes ago [-]
That variant would likely not even compile?
aarmot 4 minutes ago [-]
You can put your entropy in comments and strings
maccam94 5 hours ago [-]
DNS cache poisoning?
patmorgan23 6 hours ago [-]
Social engineering?
bityard 3 hours ago [-]
In agreement that this is good old fashioned cargo-cult security theatre, but tom7 also coined a more catchy phrase for this, he calls it "toxic max-security." http://tom7.org/httpv/httpv.pdf
kpcyrd 9 hours ago [-]
Basing your cryptographic advice on a 20 year old opinion-piece from somebody with no background in cryptography is not the flex you think it is.
wavemode 6 hours ago [-]
Appealing to lack-of-authority without actually explaining in what way his argument is wrong is significantly worse.
throwawayffffas 3 hours ago [-]
It's not cryptographic advice from an opinion piece, it's a statement about intent from the creator of the software in question.
PunchyHamster 6 hours ago [-]
But the Linus piece is sound
... for Linux
... and developers working for it constantly
the attack wouldn't work. Joe Schmoe? It's worse than just "being compromised"
You have repo of dependency locally, let's assume you downloaded good copy, the commits get compromised, you're safe.... right ?
Nope, if there is build server along the way and ESPECIALLY if it practices building from clean state every time, the build might be infected while your local copy is clean, giving no chance to notice it, unless your entire chain including local builds are reproductible AND you actually check it
kazinator 11 hours ago [-]
The problem of a SH1 collision happening by coincidence is vanishingly low and theoretical.
Nothing else matters.
Git hashes are not supposed to be a security mechanism. If your basis for trusting that you have the right checkout is the git hash, in a situation where you have legitimate concern about untrusted parties manipulating remote repositories, then you're simply wrong.
zygentoma 10 hours ago [-]
Sorry, no.
When I check out code from a git repository in a pipeline using a git hash, I expect the code to be exactly what has been reviewed by me under that hash.
Everything else would just be a crazy invitation to make supply chain attacks uncircumventable.
kazinator 9 hours ago [-]
And so if you don't trust the server that is hosted on or the security of the transport mechanism like TLS/SSL, such that the content may be manipulated by adversaries, you think that git hashes are good enough?
Well, what about someone who is fetching the commit from that server for the first time and has nothing to compare the hash against?
Oh, that would never be a problem for widely disseminated, popular, open source project, so it doesn't matter.
kstrauser 3 hours ago [-]
Rumor has it that GitHub has a flat namespace for commits. They don't store "user1/repo1/abcd1234" in one file and "user2/repo2/abcd1234" in another. Both references point to the same commit in a global shared space. If the hashes are truly unique, then that never matters, because the odds are approximately 0.000000000...000 of you and I accidentally generating the same commit. However, if I see that you pushed commit abcd134, and then I can build and push a colliding commit, and the backend doesn't check uniqueness before writes because the odds are infinitesimal that it'd ever matter, than voila, I've updated your repo by writing to my own.
Or if first writer wins, and I know that you have a popular non-GitHub repo that you're about to migrate into it, then I could pre-poison the namespace by writing my own version of a commit that I see you already have in Codeberg or Savannah or wherever.
I don't swear that this is how GitHub actually works, but I've had knowledgeable friends swear up and down that it is. And honestly, it'd make sense. They could shard storage by the first 4 digits of the hash or something, and that'd be vastly more efficient if all commits were writing to the same space.
hedora 2 hours ago [-]
I'd expect an LLM to prove this (if true) in ~ 60 minutes, given just your post and "try to prove this true or false; here's a github PAT".
That's per-repo. The claim is that user/repo/commit is flattened to just the commit. Your counterpoint (which I agree with and is well-known) is that it's flattened to repo/commit (dropping the user).
nulld3v 15 minutes ago [-]
Ah right, though I think it is rather unlikely Github flattens across multiple repos given that Github has already blogged about how their repo storage backend (Spokes) works, and it seems to operate entirely on whole repositories.
1 hours ago [-]
crote 9 hours ago [-]
Dealing with potentially-hostile hosts is quite common, actually. See for example how most Linux mirrors work, or Subresource Integrity with HTML.
Turns out securing a service to transfer a single hash is a lot easier than securing a service to transfer gigabytes of data.
Even if I don't fully trust Github, it is still incredibly convenient to be able to upload my code there and then send someone an email telling them to fetch commit `123abc` from some repo link. As long as my email isn't compromised, that should be secure.
hedora 2 hours ago [-]
Similarly, if you push to github, then trigger CI through something other than github actions, then (other than the SHA-1 problem), you have reasonable assurances that someone who has compromised GH cannot compromise your CI host.
Note that the US CLOUD Act means that, if someone figures out how to actually use collisions to compromise that CI machine, then, if the US government asks Microsoft to do use that vector to break into an overseas machine, then Microsoft will be legally obligated to do it.
PunchyHamster 6 hours ago [-]
I mean, Git commit signing should be used more often... then you can actually trust the person signing, not the distribution method
But you still need SHA256 for that
someonebaggy 1 hours ago [-]
You could pull a malicious colliding PR and reject it. Then you pull the other half of the collision without realising it is, and it's something good and you merge it. But your CI server already saw the malicious one and thinks it's the same, so builds the malicious code
shakow 11 hours ago [-]
> Git hashes are not supposed to be a security mechanism
Probably a naive question, but why not kill two birds with one stone if it can be done for a reasonable cost?
kazinator 11 hours ago [-]
Because you're not killling two birds; you're not killing the security bird with a better content hash.
A SHA-256 sum, though very good, only assures you with great confidence that you're looking at the same thing you looked at before, or that someone else is looking at elsewhere.
It is not a digital signature, and we don't want digital signatures to serve the role of content hashes.
Speaking of signatures, we have support for them in Git; you can use gpg to sign commits, and set it up to be done automatically.
Nobody is going to fake your commit such that the fake has the same SH-1 hash and your GPG signature.
The worry there is that the key holder (whether the legitimate one, or a malicious party who got a hold of the key) somehow does this: creates a new commit, signed with their key, which somehow has the same SH-1 as an existing signed commit. The git hash includes the GPG signature, so there is a significant layer of difficulty there which is likely harder than faking an unsigned SHA-256 commit.
ramses0 4 hours ago [-]
Dude... please bow out gracefully...
The attack is I pre-author `Makefile => foo: echo "hello"; bar: echo "world"` along with `Makefile => foo: echo "hello"; bar: rm -rf / ; /* $ELDRITCH_SHA1_SPIRITS_GO_HERE */` that both hash to `ff1234...`
I then prepopulate the repo with `echo "hello"`, wait 6-9 months, then submit a commit for `echo "hello" ; echo "world"` and keep (in my back pocket) the alternate implementation that also includes $ELDRITCH_SPIRITS to force a collision and MY predetermined change in functionality.
I then have free choice as to whether I serve them "hello world" or "hello && rm -rf", and THAT's the plausible problem to avoid: the ability to "cloak" content anywhere within the repo if you have enough $ELDRITCH_SPIRITS and GPU's.
You have _really_ good points, but are woefully confused. The proper answer is (would have been) to include `tree ff12354...` along with `tree-sha256 abc123456789...` for another 20 years along with a `[git.hash_strictness]: default/lazy/strict`, and some oddball `git-rerere` type packfile extension which lets you map `sha1:ff1234... => sha256:abc123456789...` "transparently" rather than the horrific situation you're laying out (correctly!) that forks the ecosystem in to "longhash" and "shorthash" when most repos don't even care in the end.
kazinator 4 hours ago [-]
> woefully confused
Yes, I didn't understand that the GPG signing just operates on the top level object in the commit and trusts the SHA-1 hashes contained in it.
The signing process doesn't recursively traverse the bytes of the commit to pull them into GPG, like you would expect.
It's like, imagine you made a "bill of materials" of your project's files consisting of their names and CRC-32 checksums, and then signed this file, and called your project securely signed, LOL.
This aspect can be fixed without foisting new hashing scheme into the content tracker. In fact, it must be fixed; users on SHA-1-based repos deserve secure signing.
It's really sneaky that the SHA-1 business (not intended to be a security mechanism) was embroiled into the signing implementation; that GPG is demoted to the strength of SHA-1.
Was that just to save some cycles? It's certainly faster just to sign the commit object!
singpolyma3 3 hours ago [-]
signatures are basically always computed over hashes. The only problem here is that the hashes are not secure. And this is being fixed.
kazinator 3 hours ago [-]
No, but, the hashes are features of the content tracking system that hook it together. There is no reason that a signing scheme must rely on and trust those hashes!
We can round up the bits that make up a commit in a SHA-1-based repo, and sign those bits securely; this is a thing that is possible.
dwohnitmok 20 minutes ago [-]
> We can round up the bits that make up a commit in a SHA-1-based repo, and sign those bits securely; this is a thing that is possible.
Yes but as my other comment explains, this is not particularly useful in and of itself.
hedora 2 hours ago [-]
Also, note that you can just sign "hello world" and "hello && rm -rf", then serve whichever you want. In the absence of a hash, then everyone has to trust you not to do that. With SHA-256 git, and a one-time audit of the code, they would need neither your public key nor to trust that the code didn't change.
kpcyrd 10 hours ago [-]
Please educate yourself what a merkle tree is. It's a well understood building block of various security systems, including certificate transparency (which explicitly uses sha256).
You refer to PGP signed Git objects, but you also argue:
> Git hashes are not supposed to be a security mechanism
Guess what the Git PGP signature is signing.
layer8 10 hours ago [-]
This is exactly right. A signature is only worth as much as the hash that it’s signing. And all the usual signature algorithms are signing a hash.
kazinator 9 hours ago [-]
The GPG signature is not signing the git hash, if that's what you mean.
The GPG signature signs some kind of hash calculated over the commit, minus the GPG header, which is thereby added.
The git hash is then calculated over the whole thing. The git hash is on the outside, and not part of the signing.
orf 7 hours ago [-]
> The GPG signature is not signing the git hash, if that's what you mean.
It kind of is - it’s signing the hash of the tree object, which is the actual thing that you’d attack with a hash collision
kazinator 7 hours ago [-]
I understand that if we sign a commit with the help of some arbitrarily strong hash, it doesn't protect the parent commit(s). The integrity of the SHA-1 hash references to the parent commits is not in question, but the authenticity of those commits themselves.
orf 7 hours ago [-]
No, not the abstract tree formed by a series of commits.
The actual git ‘tree’ object, which is the thing a commit actually points to, referenced by a hash in the commit. That is signed by the GPG signature.
crote 9 hours ago [-]
That doesn't make a difference: with sha1 a malicious change in content will still result in the same content hash, so the signature will still be valid, and the commit hash will still be the same.
kazinator 9 hours ago [-]
Only if the GPG signing process stupidly relies on the SHA-1 hash. I.e. if it takes an unsigned commit and signs only its SHA-1 hash and then creates a new commit with GPG headers. If that's how it works, that is massively stupid and can be fixed without forcing SHA-256 as a git hash. Just have the signing calculate its own digest for its own purposes.
That digest can be the SHA-256; since the infrastructure is there for it, signing should use SHA-256 regardless of what hash is used by the repository for identifying and linking content.
semiquaver 5 hours ago [-]
> if it takes an unsigned commit and signs only its SHA-1 hash and then creates a new commit with GPG headers
It does indeed. The bytes passed to GPG when constructing a signed commit look something like:
tree eebfed94e75e7760540d1485c740902590a00332
parent 04b871796dc0420f8e7561a895b52484b701d51a
author Alice <alice@example.com> 1465981137 +0000
committer Alice <alice@example.com> 1465981137 +0000
Headline
Message
where the contents being signed are entirely represented by the oids of the tree object and parent commit object. This string is very similar to the content that is fed to the hash function to produce a normal git commit object id.
kazinator 4 hours ago [-]
Haha, well that is a screw up. The weak tree hash can be attacked, replacing the content that is itself not pulled into GPG.
The "bytes passed to GPG" of course get hashed by GPG, using something better than SHA-1.
All bytes that comprise the commit should be hashed by GPG, rather than depending on the content referencing hash in the object tracking system.
This is something that is possible; it is not a logically deductive necessity that we just scan the topmost object and trust the hashes it contains.
dwohnitmok 4 hours ago [-]
> it is not a logically deductive necessity that we just scan the topmost object and trust the hashes it contains.
It kind of is. Otherwise the whole idea of signing a commit with a backing git history (rather than just a snapshot of a working directory) collapses. The only guarantee you have that the git history is what is claimed by the cryptographic signature is some sort of Merkle tree structure. Either the original one, or you have to construct a whole new parallel one with a better hash, in which case, as I bring up in a cousin comment, why not just use a better hash in your original one?
kazinator 3 hours ago [-]
In a git repo based on SHA-1 hashes, it would be good enough that the meaning of a commit signature is that the content of the commit is attested, not the parents.
Parent commits can have their own signatures, and it's true that there is an attack possible there where the same parent hash could point to two different commits, that have valid signatures of some kind (possibly from the same sneaky developer who is a bona fide project member).
Be that as it may, it's perfectly okay for the signature on a banana to validate only the banana, and not the gorilla that is holding it, or the vine the gorilla is swinging on, and the whole jungle.
It would be worth it to have better quality commit signing for SHA-1-based repos.
dwohnitmok 22 minutes ago [-]
> it would be good enough that the meaning of a commit signature is that the content of the commit is attested, not the parents.
This is a significant degradation of the implicit guarantees given by a cryptographic signature, to the point that basically all personal use cases I have for signed commits would be invalidated.
Keep in mind that git does not have diffs as first-class objects. Every commit is just a snapshot of some state of the working directory. That means that without attesting to the integrity of the parents of a commit, the only thing a commit X signed by a person A says is "at some point on A's computer, the state of the repo looked like X".
Almost all the relevant questions I would want to ask are not answered by this. E.g. there is a malicious function F that is present in X. Did A write it? Don't know. Did someone else write it? Can't know for certain. Who introduced a certain feature? Don't know. Did A sign off on a new bugfix? Don't know.
All you know is that at some point the codebase looked like X on A's computer. A might not have made any relevant changes at all!
You can only back out a diff and therefore actually attribute a change to someone (either explicitly through `git blame` or informally by looking at git logs) if you have attestation of the parents.
The Merkle tree structure of git repos is interwoven through basically ever useful thing git does. Without cryptographic signatures implicitly carrying a promise of validity for that structure, this would make commit signing useless (depending on just how broken SHA-1 is) for the needs of any org I've ever worked at.
dwohnitmok 4 hours ago [-]
> If that's how it works, that is massively stupid and can be fixed without forcing SHA-256 as a git hash.
I don't think it's massively stupid. Unless you want to re-hash the entire Merkle tree structure to sign your commit, you basically have to trust the hashes in the Merkle tree (or have a separate parallel Merkle tree) at some point in what you sign, which means you do have to trust the SHA-1 hashes. Otherwise even with a cryptographic signature you can always spoof at least the git repo history (e.g. even if you try to directly hash the entire contents of the current commit).
Re-hashing the entire Merkle tree structure seems prohibitively expensive to generate (even with a lot of caching) and pretty complicated for e.g. verifying a signature. Or you can do that incrementally, but then you're just generating a whole new parallel Merkle tree structure.
Regardless, at the end of the day, you need to trust the integrity of the Merkle tree structure. And you can either do that by trusting the hashes of the current Merkle tree, or you have to completely recreate a new one with more trustworthy hashes, in which case why not just use better hashes in your original tree?
kazinator 3 hours ago [-]
I need a commit signature to only attest the integrity and authenticity of that commit, nothing else; no parent or grandparent, etc.
pseudohadamard 9 minutes ago [-]
That's what a lot of the responses are missing, which is why the response to them in turn is both no and yes. There's no published threat model that I know of for the properties that the hashes are supposed to be providing for git, which means anyone can make up any required property they like and then confidently state that SHA-1 provides or does not provide it, see this discussion for examples. The result is, to use another Linus quote as the OP has already quoted him in the writeup, "people wanking around with their opinions".
We use SHA-1 in our storage mechanism, which predates git. There is a (quite long) written threat model. Someone being able to generate collisions with an enormous amount of effort under just the right conditions is not a threat under that model.
SAI_Peregrinus 8 hours ago [-]
> Git hashes are not supposed to be a security mechanism.
Commit signing indicates otherwise.
4 hours ago [-]
throwawayffffas 3 hours ago [-]
The important point is that the switch is breaking backwards compatibility. The proposed solution, a new independent hash just for verification makes sense.
plorkyeran 1 hours ago [-]
Do I understand correctly that GitHub currently doesn't support SHA-256 repos at all, so the author faked a screenshot of what the create repository UI would look like if they choose a really stupid way to add support for SHA-256 repos, and then proclaimed it an unsolveable problem? Selecting the repository format based on the first thing pushed to it is really not a crazy idea.
The submodule problem is real, but the forge problem is entirely just "if forges implement support in a way that makes it painful it will be painful", and that's true of literally every feature that requires forge support.
larusso 1 hours ago [-]
I also don’t know how big this problem really is. You practically have two ways of creating a repo. Local first and push or create a repo remote with Readme etc and clone.
The submodule issue is a different beast. But that one is currently a problem as well when a person choose a http url or ssh url. I have some extra git configs to normalize everything to ssh for instance.
gandreani 11 hours ago [-]
One of my favorite fun facts about Fossil SCM (another source control by the devs of sqlite) is that they patched their use of SHA1 6 days after the shattered attack was published:
"Both Fossil and Git started out using only SHA1 hashes. But when the SHAttered attack against SHA1 was published on 2017-02-23, the need to migrate to a stronger hash algorithm was recognized. Fossil added the ability to use SHA3-256 as an alternative on 2017-03-01 (six days after the SHAttered attack was first published). SHA3-256 is now the default for all new repositories and check-ins in Fossil, though older check-ins that occurred prior to SHAttered can still use their original SHA1 hash. Hence, no repositories had to be rebuilt and no hyperlinks were broken."
To me it's so interesting watching in realtime Git is still battling with this decision and for Fossil it was just another week of development.
That whole page is fun to read. Another fun fact somewhere else in the docs is that Fossil uses a grow-only set to store commits. They came up with this scheme some years before it was formalized by CRDTs!
bawolff 3 hours ago [-]
in fairness, git switched to sha1-DC in may 2017, so they were only a few months late in mitigating it.
schacon 11 hours ago [-]
I mean, there are two things here. One is how difficult it is to have a different hashing mechanism. Brian and other heroes in the Git core group have done amazing work to make this _technically_ possible on a repo level. To test some of my theories, I trivially implemented MD5 and an insanely dumb and easily breakable hash backend. It's not _hard_ to change the mechanism now. It's about the community.
Fossil isn't difficult to change not because it's technically harder for Git but because Git has a community and ecosystem that Fossil does not. The cost is not in the individual project for Git, the cost is because there is _so much_ in Git and this bifurcates everything.
schacon 11 hours ago [-]
Also, interestingly, Git today does _not_ use a straight SHA1 because of these attacks. It uses `sha1dc`, a slower collision detecting variant that specifically checks for this vector of attacks. So currently, Git's SHA-1 variant is not susceptible to the SHAttered/Shambles attacks.
gandreani 11 hours ago [-]
Agreed! This isn't a tech dig at all.
To me it's more of a reality of creating a tool with a huge active community and a community of contributors and creating a tool with a small team and small community.
6thbit 11 hours ago [-]
That's impressive. I suppose they had a more flexible architecture to make that change so fast.
Is there any writeup on why it was easy for them and not for git?
toymin 11 hours ago [-]
My guess is that it's less about the architecture and more about the blast radius and the number of users
gandreani 11 hours ago [-]
Hmm probably nothing architecture wise. It's probably just the fact that fossil is developed by way fewer devs.
From the skim I read of this article it seems both projects arrived at the same solution: support both but make SHA-256 the default.
kccqzy 4 hours ago [-]
No it’s different. Git supports both but you cannot mix them in a repo. Fossil allows mixing them in a single repo. This is clearly stated in the link.
rurban 3 hours ago [-]
SHA3-256 != SHA-256
hedora 2 hours ago [-]
The saga continues with that proposed terrible github UI in the article (from a github cofounder!)
fragmede 11 hours ago [-]
It's easier to make world breaking changes when the world is really small. If git could magically just get everything and everyone to cut over and use git 3.0 in a magic instant, it wouldn't be having this problem.
samus 20 minutes ago [-]
It seems Fossil can use multiple hash algorithm within the same repository. Yes, that decreases overall security since the hashes pointing to older objects can still be spoofed.
nofunsir 11 hours ago [-]
Here's a stoichiometric bird for you:
:%s/git 3\.0/python 3.0/g
fragmede 10 hours ago [-]
Or IPv6 or windows 11.
:x
Neovim's lazyvim plugin sucks because it takes over H C and L.
0x00cl 3 hours ago [-]
I think this change is more to do with politics rather than "security". Those kind of things where companies or gov, need to be certified with those super secure certificates and can't be using software that uses SHA-1. I don't have proof, but I'm not doubting it either.
This is what I saw in one of the mails.
> > There are organizations where SHA-1 is blanket banned across the board - regardless of its use
And also on git 3.0 breaking changes.
> > SHA-1 ... recommended against in FIPS 140-2 and similar certifications
Since SHA-1 isn't used for security in git, they should've instead moved to a non-cryptographic hash function such as MurmurHash3 and avoid all these problems, instead of moving to SHA-256 until SHA-256 is broken and need to move to the next cryptographic hash that is now incompatible with previous versions of git repositories.
hedora 3 hours ago [-]
SHA-1 is used for security in git. It's the thing that guarantees a commit SHA is unique. Without that, you open up all sorts of downstream infrastructure to supply chain attacks, where old objects get replaced with malicious ones, and then replicated on each subsequent git pull.
Linus' old argument was that the substitution would probably be noticed eventually, but that's specific to the way Linux uses git, and what he said probably isn't true in practice -- even if it is, there have been enough supply chain attacks since then to prove that even temporarily serving the wrong stuff to developers or CI is enough to allow lateral movement into other packages, production machines, etc, etc..
> There are organizations where SHA-1 is blanket banned across the board
This is very likely the case. And if it is, then it's a lost battle. You simply can't reason with that kind of corporate people, let alone have an argument around this level of complexity. Kafka (the writer, not the message broker) predicted this 100 years ago.
When going through the article, my instinct was changing from "annoying" to "this really sounds like a Python 2/3 moment for Git" to finally "oof this is going to be a mess" in the libraries/submodules part.
meinersbur 11 hours ago [-]
Linus Torvalds in 2007:
> but the point is the SHA-1, as far as Git is concerned, isn't even a security feature. It's purely a consistency check. The security parts are elsewhere, so a lot of people assume that since Git uses SHA-1 and SHA-1 is used for cryptographically secure stuff, they think that, Okay, it's a huge security feature. It has nothing at all to do with security, it's just the best hash you can get. ... [1]
So Torvalds used SHA-1 purely because he needed a hash function with no other property than identifying content.
creata 3 hours ago [-]
> It's purely a consistency check. The security parts are elsewhere
Sorry if the video answers this, but how does commit signing work if it doesn't rely on the hash algorithm being resistant to at least second-preimage attacks?
someonebaggy 1 hours ago [-]
Things changed since 2007
zamalek 11 hours ago [-]
Exactly. But that's why I think SHA was a mistake. He should have gone with something like murmur to avoid all this frothing at the mouth.
layer8 10 hours ago [-]
It was a mistake to assume a fixed algorithm in the repository format and client-server protocol. I remember being surprised when I learned about that choice, being familiar with cryptographic protocols and formats where the hash algorithm is usually a parameter that can vary for each concrete hash.
loeg 6 hours ago [-]
> being familiar with cryptographic protocols and formats where the hash algorithm is usually a parameter that can vary for each concrete hash.
This flexibility ("agility") in cryptographic protocols is often seen as a mistake today, actually.
WorldMaker 4 hours ago [-]
Cryptographic protocols have been moving towards something of a compromise in flexibility. "Everything flexible" is a security risk, especially when "everything" includes "fallback to nothing secure". "No flexibility" is a security risk because you can't upgrade. The middle path is something like "version numbers" with hard breakpoints. "I only support v2 of this cryptographic protocol and will not fallback to v1."
Which is sort of the hash algorithm approach git is taking with incompatible versions and a version break.
throw0101c 9 hours ago [-]
> It was a mistake to assume a fixed algorithm in the repository format and client-server protocol.
See also perhaps Wireguard, which touts itself as not having "cryptographic agility" because they wanted to avoid all (perceived) problems and complications of IPsec. But now that PQC is (allegedly) approaching there's no easy to update things because (AIUI) there's no negotiation possible in the protocol; you're basically standing up a 'Wireguard 2.0' that runs separately than the original.
computerfriend 50 minutes ago [-]
Wireguard is secure against a quantum computer though, via an additional pre-shared key.
> If an additional layer of symmetric-key crypto is required (for, say, post-quantum resistance), WireGuard also supports an optional pre-shared key that is mixed into the public key cryptography.
Which is fine, for wireguard which only encrypts things ephemerally.
akerl_ 4 hours ago [-]
That is in fact the idea, and was on purpose.
throwawayffffas 3 hours ago [-]
1. As others have noted, flexibility in cryptographic protocols is generally a mistake.
2. The hash function is not used for cryptographic purposes!
Someone 10 hours ago [-]
Linus, in 2005, couldn’t have gone for murmur, from 2008.
Was there “something like murmur” in 2005 that’s cryptographically better than SHA1?
sgerenser 7 hours ago [-]
Yeah, in 2005, SHA-1 was just about the best you can do given the constraints of the time (without picking something much more esoteric, much slower, etc.). Using SHA-256 at that time would have been noticeably slower on the computers of the time, and made repo metadata take up a lot more space.
bawolff 11 hours ago [-]
If that was true, i doubt git would have switched to using the slower version of sha-1 that detects attacks.
huflungdung 11 hours ago [-]
[dead]
kazinator 11 hours ago [-]
Frothers gonna froth, though.
UltraSane 5 hours ago [-]
Hard coding a specific hash algorithm was a big mistake.
Notably, a few of the featured author's reservations appear to be addressed. According to the Git docs:
- Objects can be referred to by their old, SHA-1 name or their new, SHA-256 name. This means old refs in docs and comments and such remain valid. The mapping between SHA-1 representations and SHA-256 representations appears to be intentionally bijective a.k.a. 1-to-1 (assuming no hash collisions), so that it could be re-computed on demand. The constraint of bijectivity appears to be the source of some limitations, ex. no mixed repos and submodules needing to match hash algroithm, but also bijectivity has strong benefits like the following items.
- A bi-directional dictionary is maitained from SHA-1 to SHA-256 names so translations between the two don't required re-hashing objects. This table could be recomputed on demand due to the bijection between names; it's only a performance optimization.
- A local SHA-256 converted repo (including an SHA-256 converted submodule) can interoperate with an SHA-1 only remote transparently to the remote server by translating names using the lookup table.
- SHA-1 based GPG signatures will be preserved. A commit can be signed based on its SHA-1 representation, its SHA-256 representation, both, or neither. The bijection means the two types of signatures are in a sense interchangeable, or in other words the bijection between object representations implies an equivalence relation on signatures. An SHA-256 converted repo can quickly validate an SHA-1 based gpg signature using the lookup table.
computerfriend 23 minutes ago [-]
The hashing algorithm's security properties are a security property of Git. This is because:
* we pin to commit hashes and expect this to refer to immutable content,
* commit and tag signatures are over the hash.
The Linux quote about trusting the distribution doesn't make sense to me, as Git is content-addressable and decentralised, although possibly at the time it was a reasonable position to take for kernel development, but [1] could equivalently happen for Git and the hash algorithm being non-broken is required for it to be noticed. Not having to trust the forge is a very desirable property.
I don't understand why Git is not making the SHA-1 and SHA-256 modes far more compatible with each other.
SHA1-hashed objects should be able to refer to SHA-256-hashed objects, although this seems somewhat pointless.
But SHA-256-hashed objects should also be able to refer to SHA1-hashed objects, with a major caveat: if those objects themselves are part of a collision pair, then there is a genuine problem. But this is avoidable! Suppose that Linux decided to migrate to SHA-256. The upstream project could choose a pair of dates, say January 1 2027 and March 1 2027. Up to the first date, maintainers would be welcome to submit hashes of objects that are not yet in the repo but that they think they might submit later on, and, on that date, the upstream tree would finalize the list of these objects and reference it in the repo (with a new mechanism for this purpose). Effective the second date, the repo would start publishing SHA-256 commits and would never again accept a SHA1-hashed object that was not in the repo at the cutoff date or referenced as part of the Jan 1 block.
And now it would be impossible to get a new SHA1 collision in to the repo.
The only new git features needed would be:
a) actual compatibility so that a SHA-256-hashed object could reference a SHA1-hashed object
b) a new object type that's a list of allowed SHA1 hashes (or probably a tree of them) that is itself hashed with SHA-256 and a mechanism to link to one of these from a commit
c) a policy mechanism to set a repo to only allow SHA1-hashed-objects that a reachable from a preconfigured SHA-256-hashed commit
schacon 12 hours ago [-]
Emily's talk does a pretty good job of summarizing the issues with intermixing the hashes: https://youtu.be/eJJp0RE7cd4
RJIb8RBYxzAMX9u 12 hours ago [-]
I skimmed the video, and I didn't quite catch that. Near the end of the video, however, she did mention that interop is in the works[0].
In any case, even if Git 3.0 were completely incompatible, it would suck, but it's not the end of the world. You just treat it as if you were migrating from one SCM system to another. CVS -> SVN -> Perforce -> Git -> Git 3.0 -> [...] been-there-done-that. This is something that both open-source and commercial projects have had to deal with over the years.
Or maybe it would be a repeat of Python 2.x -> 3.x. ¯\_(ツ)_/¯ With AI assistance, hopefully porting the tooling over may go a lot quicker and smoother.
Her entire section 2 (starting at 6:12) is basically about why git does not and will never allow mixing of SHA1 and SHA256.
The interop discussed is using copybara as a copy tool to move data from SHA1 based repos to SHA256 based repos and vice versa.
RJIb8RBYxzAMX9u 11 hours ago [-]
Thanks. I went back and re-watched that part, and her point's that "[a] tree's cryptographic strength is equal to the weakest hash algorithm anywhere in the tree," and that's fair. I also understand why Git 3.0 may not want to give user the choice, though I would still rather it be given.
amluto 10 hours ago [-]
> "[a] tree's cryptographic strength is equal to the weakest hash algorithm anywhere in the tree"
This is simply wrong IMO, for two reasons:
1. The attack on SHA-1 is a collision attack. Once you have frozen a hash, you cannot attack it with existing cryptanalysis. If there were a preimage attack it would be a different story.
2. Even if there were preimage attacks, one could freeze a mapping from SHA1 hash to SHA-256 hash.
In fact, #2 seems like en excellent design. Objects could reference such a mapping, and a repo could disallow conflicting mappings (the mappings would only be accepted if the mapped objects are reachable from the mapping and the mapping is correct).
schacon 11 hours ago [-]
They do give the user the choice, but the default is changing. My point is not necessarily to rip out the SHA-256 option, but simply to not make it the default. Because then people will create repos in that format that do not understand the ramifications, where the opposite should be true.
schacon 11 hours ago [-]
My point is not that it's the end of the world (or the end of Git), but that it will be painful and unclear and confusing to lots of people. That would be fine if it made a huge difference in trust or protection, but it's the wrong way to do that.
mort96 11 hours ago [-]
Hm but the date is stored inside of the commit. The only way we can know that a commit's date is authentic... is through its hash. If I can forge commits with any SHA1 hash at will, I can make a repository whose head commit has the same SHA1 as the one in torvalds:
/linux but where any commit was replaced by a malicious commit with the same SHA1 and a fake date. You have no way to detect that my repo is inauthentic other than through a deep history comparison. The whole idea behind a merkle tree is that just checking the hash of the top is sufficient to know the identity of the whole tree.
I don't know what the solution is, but I'm inclined to believe that any repo with a single SHA1 commit is as weak as a repo with all SHA1 commits.
amluto 10 hours ago [-]
The date that a repo receives a commit is known to that repo. And a repo can stop accepting new SHA1 objects. And a SHA256 object could have a flag that says that no SHA1 objects may ever reference it.
mort96 9 hours ago [-]
The design of Git, as a Merkle tree, is meant to allow for use cases like this:
* I host a mirror of the Linux git repo.
* You download Linux from my mirror.
* You check out a commit, say fd179f8a05be3ccae366b9b96e176b51fbe54aab, which you know is a genuine commit through some out-of-band mechanism (mailing list, GitHub web interface, a line in a Nix file, whatever).
* You check whether the repository I gave you is legitimate or not by re-computing the hash of the commit which I claimed was fd179f8a05be3ccae366b9b96e176b51fbe54aab. If it comes out to be fd179f8a05be3ccae366b9b96e176b51fbe54aab, you know it's legitimate. If it doesn't, you know it's fake.
This is a completely normal use of Git. People download from mirrors all the time. People rely on commit hashes to identify a specific source tree. People trust that if whatever the mirror gave them hashes to the right value, it's genuine. That way, you don't have to trust the mirror.
If I can forge my own commits to have any hash I want, this whole model breaks down. I can replace some old commit in the repo with my own forged commit with the same hash, and when you download a copy of the Linux repo from my mirror, you'll receive a repo with malicious content, but it'll hash to the same fd179f8a05be3ccae366b9b96e176b51fbe54aab hash as a genuine repo would. This breaks the security model of Git.
amluto 9 hours ago [-]
> You check out a commit, say fd179f8a05be3ccae366b9b96e176b51fbe54aab, which you know is a genuine commit through some out-of-band mechanism
That's a 160 bit hash, which is SHA-1, which has the security properties of SHA-1.
Suppose you check out a commit with a given SHA-256 hash. That commit object represent the root of a tree where all the edges are hashes (and types, etc). I'm suggesting one of two designs:
a) (Simpler but weaker) If Linus has published that commit, then he is confident that he hasn't pulled in any too-new SHA-1 hashes and that there are no collisions present in what he thinks the tree is. So, by induction on the traversal depth, there is only one actual object identified by each edge, and those objects contain the hashes of their child edges, so those hashes are all correct.
This breaks if there is a malicious collision already in the tree.
b) (Stronger but higher overhead and more complex) There would be an object or objects, discoverable from the root by following only SHA-256 edges, that encode a duplicate-free mapping from SHA-1 hash to SHA-256 hash. The client finds and parses that and then, as it traverses the tree, each time it reads a SHA-1 hash, it computes the SHA-1 and SHA-256 hash of the referenced object, verifies that the pair is in the mapping and also verifies that the SHA-1 hash matches what the edge requires.
I think that (b) is genuinely cryptographically secure in the sense that, if you can construct a commit that has the same SHA-256 hash as an official upstream commit but different contents, then there is necessarily a SHA-256 collision.
mort96 8 hours ago [-]
For A), I don't understand what the point is? I never mentioned what Linus is confident about, I talked about what you can verify when you pull from my mirror. I could replace a commit from 2010 with a malicious one
For B), I would think this could work, but it's a completely different solution from what you proposed and what I responded to.
amluto 6 hours ago [-]
> I could replace a commit from 2010 with a malicious one
How? Remember, there are (currently, anyway) no known SHA-1 preimage attacks.
> If I can forge commits with any SHA1 hash at will
We probably don't want to wait until there are practical pre-image attacks discovered to change away from SHA-1.
amluto 6 hours ago [-]
This is a fair point.
I think I stand by my second proposal. I also think it's absurd that, after all these years, upstream git still can't figure out a credible migration plan.
PunchyHamster 6 hours ago [-]
If there is one way enforcement (i.e. there is one point where the last SHA1 commit was signed by first SHA256 commit), I think it should be safe ?
The "commit before" might be compromised, but the git commits refer a snapshot of a tree + a list of previous commit IDs, so the "new" SHA256 commit will not have any files altered
watinthedeutsch 39 minutes ago [-]
> The first thing that you'll notice (other than the much longer hash value) is that you can't push this code to GitHub, though that will almost certainly be fixed by the time Git 3.0 is released. In fact, that's probably the main thing currently delaying 3.0 entirely.
> But when you do want to push it to GitHub (or any host), you will need to tell them when creating the repository on the server that this is a sha256 project. Every project will now be in one bucket or the other and they cannot be mixed.
> This is immediately going to frustrate people because now they need to know what version of git they ran git init with and make sure when they go to GitHub to create the server repository, they choose the right one.
Github can fix this problem, and we have had siilar stuff in the past, when we had to switch repos. This is simple and doable, just need gradual work. I don't see the problem.
gorgoiler 47 minutes ago [-]
This article says if you manually enable the experimental new code today then “you can't push code to GitHub” and that old and new repos “cannot be mixed”.
The git transition document talks about keeping a bidirectional mapping between old and new hashes, and converting between the two when pushing to legacy repos:
I suspect that’s what gitbutler.com means when they say that, although broken today, everything “will almost certainly be fixed” when git 3.0 ships.
While I’m sure they make a good argument in the hash theory section, I’m less inclined to believe the “train wreck” part of their post.
If you upgrade your house and leave the roof off then yes, it will be a “costly mistake” the next time it rains, but if your plans say you intend to put the roof back on then I’m not sure how I am supposed to interpret a blog post warning of the perils of a roofless house.
nicoburns 12 hours ago [-]
From what I'd read, SHA256 in git is showing every sign of being another IPv6. In particular:
- It's implemented in a non-backwards-compatible way
- The benefits over the older model are a bit nebulous
- There's a large amount of tooling that needs to catch up, and little sign that there is movement there
kccqzy 3 hours ago [-]
It could also be another Python 3 situation: backwards incompatible, unclear benefits with many downsides (3.0 and 3.1 being very slow), large number of libraries that need to catch up.
sltkr 12 hours ago [-]
The difference with IPv6 adoption is that the internet relies heavily on network effects: so long as some hosts only have an IPv4 address, you need an IPv4 address for full connectivity, but then if everyone has an IPv4 address anyway, there is no immediate need to migrate to IPv6.
(Yes us Hacker News users have plenty of use cases for IPv6, like self-hosting and peer-to-peer networking and so on; we are not the average user.)
This effect doesn't exist for the Git migration. Each repo can be updated independently; it doesn't affect users of other repositories, and most likely, the majority of devs will work on some SHA-1 repos and some SHA-256 repos with no issue.
If anything, I would compare it with the Python 2 to Python 3 migration, which was also painful, but succeeded eventually (despite being much less necessary in the first place).
kllrnohj 6 hours ago [-]
> Yes us Hacker News users have plenty of use cases for IPv6, like self-hosting
Funnily enough, self hosting is why I can't use IPv6. I want vlan isolation, but only get a /64 from my ISP.
Fortunately the lack of IPv6 also isn't a meaningful loss anyway so whatever
charlieyu1 6 hours ago [-]
I just started learning IPv6 with AWS since they charge $0.005/hr per IPv4. Maybe it will be more expensive in the future and eventually it will be the new default.
metalliqaz 11 hours ago [-]
changing repos to the new IDs would break any existing links to content on the pre-migration repos.
AndrewDucker 10 hours ago [-]
Unless you link using tags.
crote 8 hours ago [-]
Which is considered a Really Bad Idea because tags aren't immutable, so there's absolutely zero guarantee that it'll point to the same commit a few months from now.
The GitHub Actions ecosystem found out the hard way, through some rather high-profile compromises. They hotfixed it by adding "immutable tags" to their platform, and are now working on adding a lockfile to... easily reference a commit hash.
PunchyHamster 6 hours ago [-]
well if you use a knife to stab your fingers that's not a knife's fault
repo can also rewrite existing commit and you again won't be able to retrieve it so switching to commit IDs only lowers the level of failure somewhat
computerfriend 44 minutes ago [-]
Having a fetch break is better than having a fetch pull down malware.
gspr 12 hours ago [-]
> - The benefits over the older model are a bit nebulous
This is far from the case with IPv6!
post-it 12 hours ago [-]
Is it? There are many benefits in principle to IPv6, but if my ISP continues to assign me a single dynamic IP, those benefits are entirely moot for me.
bityard 12 hours ago [-]
If your ISP is assigning you a single IPv6 address, they are doing IPv6 wrong. You should be getting your own /56.
ghusto 11 hours ago [-]
Every time you use the word "should" you're fighting reality.
bigstrat2003 5 hours ago [-]
If an ISP chooses to deploy a technology in a bad way, that is not the technology's fault. Even if most ISPs were just giving users a /128 address (which to be clear, is very much not the case), that would not be some kind of failure of IPv6.
gspr 9 hours ago [-]
Plenty of ISPs do it right (examples: my last 3). That the parent poster's apparently doesn't isn't an argument against the technology their ISP fails to deploy correctly.
Reality is that lots of ISPs do this right. And have for a very long time.
bigstrat2003 12 hours ago [-]
Even a /64, while wrong, would be a significant improvement over IPv4.
mort96 12 hours ago [-]
Your ISP assigns you a single dynamic /128? Which ISP is that?
post-it 10 hours ago [-]
My ISP (Bell Aliant) doesn't support IPv6 at all. Rogers though, for example, assigns you a /56 block but it's dynamic.
thenewnewguy 12 hours ago [-]
Is the argument that there are no benefits to IPv6 for anyone/society because your specific ISP messes it up?
post-it 10 hours ago [-]
Of course not. The argument is that the benefits are nebulous.
gspr 9 hours ago [-]
... if you use ISPs that don't deploy the technology correctly. Faced with multiple providers of a technology, are we to judge its benefits by the worst ones?
embedding-shape 12 hours ago [-]
Right, but just because you don't happen to have IPv6 right now, how does that remove the benefits for others to have IPv6? That's like saying having a faster CPU wouldn't mean faster performance, because I don't have that CPU yet.
12 hours ago [-]
Onavo 12 hours ago [-]
There's a massive push right now from top down to have secure software supply chains. Google SBOM and SigStore. It's not an organic need but if you have government customers you don't have many options.
Dayshine 12 hours ago [-]
Ironically rewriting git history is a perfect opportunity for a supply chain attack.
valmyr 2 hours ago [-]
This is actually a good change. If you want to change the security assumptions of Github repository, ie make them somewhat distributed. Then the SHA-1 based commit hash is a major problem. It only costs about 10k in 2024 to find a collision to a random SHA-1 hash. While this costs essentially makes the attack infeasible for most threat models. It does limit how far you can scale this without an obvious footgun waiting for you.
This is a good change, even though there is a massive technical debt in changing such a widespread system. It is worth the effort. Should generations from now still be using SHA-1 for their Git ops? Sometime you have to do the switch, otherwise you will never progress.
For my usecase basically i needed to know that every Git commit pointed at a cannoical blob. With SHA-1 you could generate two blobs which hash to the same SHA-1 hash, while you can do the format verification which helps i could not do that in my usecase as i did not know the underlying data. To fix this i had to very ugly have two methods of referencing any Git object, a cryptographically secure SHA-256 ID and the Git ID SHA-1.
purpleidea 12 hours ago [-]
This means, if you migrate your repo, every single commit message that contains text like: "please see commit <sha1>" will now be broken.
This will be a train wreck. I hope they don't release before adding compatibility modes to keep the existing sha1's around in the database.
OneDeuxTriSeiGo 3 hours ago [-]
Note that there exists multiple ways to continue to lookup SHA1s in a SHA256 repo.
One example is to maintain git-replace refs for the rewritten SHAs but there also exist config flags to enable object format compatibility extensions that help translate the SHAs back and forth.
Tools like git-filter-repo[1] support rewriting commit hashes in commit messages. git-filter-repo actually does it by default; see `--preserve-commit-hashes` in the manual[2].
I need git-filter-repo to rewrite entire documentation and also resurrect and rehire earlier employees to repeat their GPG signatures.
donatj 11 hours ago [-]
Sure, but that's not going to rewrite Slack messages, emails, GitHub links, docs
purpleidea 11 hours ago [-]
Came to say something exactly like this... The old sha1 handle needs to be still available in the same way an HTTP 301 redirect would work.
diegocg 11 hours ago [-]
... do not migrate old repos? I'm not sure why people would do that. Or, if they do, why would they replace the current repo name instead of creating a different one and keeping the old one closed to make the references work.
I don't think this is going to be a problem at all.
krupan 1 hours ago [-]
If we don't need to migrate repos to sha256 them why do we need to change git to use sha256?
beart 9 hours ago [-]
Hmm. If this is a security issue that matters, would you not be expected to migrate?
Not talking about archived code, but active projects that pre-date git 3.0
112233 11 hours ago [-]
Once this starts being actual pain, we will each vibe the replacement index creator (git already supports replacement objects), for back-forth conversion, populated on pack and object indexing.
For massive perf and mem use damage. But oh well. And then we will wait for official version
windsurfer 11 hours ago [-]
Since SHA-1 is already broken (just expensive in terms of GPU-time), then the text "please see commit <sha1>" is also already broken.
Dylan16807 11 hours ago [-]
You can't attack an existing normal commit.
But also collisions there aren't a big deal. People will cite short hashes when referring to things and that's not "broken".
windsurfer 8 hours ago [-]
If it's an existing commit and you're already converting the repo, you can just convert the commit messages as well.
schacon 11 hours ago [-]
That is essentially only a second preimage problem, which is basically impossible.
10 hours ago [-]
schacon 11 hours ago [-]
There are plans to keep sha1s around in a database, but as far as I know, no way to transmit those, so they seem specific to individual forges. They can be recomputed, sure, but again, any signatures break and it's possible that in the case of an actual replacement, the recomputation is now wrong and not easily comparable. So what is the point?
purpleidea 9 hours ago [-]
This is not about forges, it is about the repo I have in a folder on my computer. The sha1 hashes shouldn't go away. Yes the forges also need to support this.
6thbit 12 hours ago [-]
I thought this would be a snark but it's an extremely well put together argument against the "Hashmageddon".
If you're replacing the weakness of SHA-1 just by going to another algorithm, you better be prepared to go to the next one when sha256 collisions happen, and it doesn't sound like git's design would be easy to modify for this type of crypto agility.
I do like their proposal for using signatures to establish trust and allow swapping sha256 for whatever comes next.
schacon 11 hours ago [-]
Technically, git's design (thanks to very smart people trying to solve this problem like brian and others) is _very_ easy to modify to different hashing algorithms now. A lot of amazing work has gone into this in recent years.
However, it's not a git problem. It's an ecosystem problem. It's that every git repo has to choose one and they're entirely incompatible with each other. That is the cost and the difficulty.
UltraSane 4 hours ago [-]
Git should support multiple hashes for commits
bawolff 11 hours ago [-]
I think there is a question though when that will happen and if it will be in our lifetime. SHA-1 started showing weakness in 2005 (collision in 2^69 instead of expected 2^80. This was later brought down to 2^61 in 2011), the same year git was invented. Nobody has found a similar weakness in SHA-256 as of yet. SHA-256 is still at its design strength of 2^128
It took 20 years to go from vulnerability in sha-1 to having to replace it out of caution. There is no such vuln in sha-256 yet. It could easily be 25 years before we find one, and another 25 years before we have to do something about it. Perhaps longer. Will git still be used 50 years from now?
6thbit 10 hours ago [-]
With the kind of compute power available nowadays and AI models I wouldn't be surprised we see it much sooner.
All it takes is just one collision to consider it broken right?
But hey maybe the attempt to fix it makes git controversial enough it falls out of favor, and nobody uses it anymore in 2 years, problem solved? sure.
bawolff 9 hours ago [-]
> All it takes is just one collision to consider it broken right?
No, its considered broken before that stage. i.e. when someone discovers an attack that would allow someone to create a collision faster than they should while still being impractical.
> With the kind of compute power available nowadays and AI models I wouldn't be surprised we see it much sooner.
Computer power doesn't super matter, what matters is algorithmic breakthroughs. So far i dont think there are any examples of major breakthroughs of that type via AI, although perhaps i am just misinformed. Its still early in the AI revolution, it might still happen, but as it stands i don't think there is any reason to worry about that.
TheRealPomax 11 hours ago [-]
Perhaps not clear enough, but Scott was a cofounder of GitHub, so he knows a thing or two about git in the real world =)
schacon 11 hours ago [-]
Being a founder of GitHub doesn't make my opinion more interesting. I hope the argument stands no matter who wrote it. :)
TheRealPomax 7 hours ago [-]
It does though, even if it shouldn't be blindly taken as gospel. Arguments help, but some folks have a better brand of apple box to stand on, and "I live and breathe git" helps quite a bit ;)
MBCook 12 hours ago [-]
So they’ve been talking about this for many years, planning, and finally announce when they’re going to switch the default.
So this is the right time to post that everything they’re doing is wrong? Did you engage in all the discussions about it and how best to handle it? Whether SHA-256 was the best solution?
I don’t see anywhere that it talks about alternate proposals or why they might have been better. Why the particular suggestions here were rejected.
This seems like a bunch of Monday morning quarterbacking.
schacon 12 hours ago [-]
I do mention this in like the first paragraph. I don't feel great about it, but I've listened to these issues for years now during contributor summits and Git Merge talks and while it's always seemed problematic, I thought they would come up with a good solution. This last Git Merge confirmed that it's close to the switch and not in any way solved or improved. I don't want to just go with it for groupthink reasons. I never thought it was a good idea and I have said that, but we have a last chance to rethink this, so I'm curious if I'm alone or in the silent majority.
throwworhtthrow 12 hours ago [-]
Your argument is persuasive and well illustrated. I think the problem is the intro paragraphs come off as too certain of catastrophe which, when juxtaposed with your claim that "smarter people than me have been working on this", makes it sound like you don't actually believe they're smarter than you. The rest of your essay feels fair and not judgmental.
schacon 11 hours ago [-]
I do believe they're smarter than me, but sometimes very smart groups talk themselves into ultimately impractical solutions because they're all smart. Sometimes you need a dumb guy to come in and say "are you sure this is right?"
throwworhtthrow 11 hours ago [-]
I believe you are sincere. But "is about to be a huge, costly, global train wreck" lacks the nuance of "are you sure this is right?" and will rub some people the wrong way.
Edit: I'm not suggesting you should have written it any differently. I think you made the right choice to be a bit provocative because it grabs the attention that's needed.
schacon 11 hours ago [-]
I do believe this. But that doesn't mean I can't be convinced otherwise by a good argument. The point of this post is to see if anyone has a great counterargument to change my mind.
nofunsir 12 hours ago [-]
The plans have been on display in a cellar. Beware of the leopard.
bawolff 11 hours ago [-]
This isn't really applicable. The plans have been talked about for a while very publicly.
MBCook 12 hours ago [-]
I don’t understand what this is supposed to mean.
NikolaNovak 12 hours ago [-]
It's a hitchhiker guide to the galaxy reference, where sure, something is technically available but not clearly published and there are hoops even for those who know what they're looking for.
(No clue if it's applicable here, I'm not aware of this case, but I believe that's the reference if it helps :)
Edit : exact quote, as Arthur's house is about to be demolished for a highway bypass:
"But the plans were on display…”
“On display? I eventually had to go down to the cellar to find them.”
“That’s the display department.”
“With a flashlight.”
“Ah, well, the lights had probably gone.”
“So had the stairs.”
“But look, you found the notice, didn’t you?”
“Yes,” said Arthur, “yes I did. It was on display in the bottom of a locked filing cabinet stuck in a disused lavatory with a sign on the door saying ‘Beware of the Leopard.
Dylan16807 11 hours ago [-]
It's not at all applicable.
zygentoma 10 hours ago [-]
I believe it's a reference to the hitchhikers guide to the galaxy – where the plans to remove the protagonists building to build a bypass road was hidden in this way.
fragmede 11 hours ago [-]
nofunsir is a stoichastic parrot, matching to a bit in Hitchhikers Guide to the Galaxy, wherin the protagonist should have known to protest a plan to demolish his home where plans where clearly documented in a hard to find place that they could not have known about. It is not a good pattern match, because git has been discussing this in public on documented mailing lists for years.
nofunsir 11 hours ago [-]
No I'm not. Real thumbs here. If anything this paragraph sounds like LLM slop
fragmede 11 hours ago [-]
Run it through pangram if you want, I wrote it with my human brain. What I did not do, however, was parrot an inapplicable quote.
In what way has git's discussion of their move been hidden away in a metaphorical basement?
pcthrowaway 5 hours ago [-]
That comment was short enough that you'd really have to be naive to think a tool could reasonably detect if it was authored by an LLM.
Additionally, as someone who read Hitchiker's guide to the Galaxy ~3 decades ago, I managed to suspect it was a reference to it, even if the reference wasn't the most applicable
nofunsir 11 hours ago [-]
I'm sure the cellar was also documented, Arthur eventually found it, did he not?
Mailing lists are basements in 2026
fragmede 11 hours ago [-]
Since 1999, we've had Google to help people find things. True, a mailing list isn't an Instagram reel delivered directly to your face with audio and blasted out on Fox News, but if one was interested in the development of git, an LLM or a Google search would readily tell you about the existence of those mailing lists.
kazinator 11 hours ago [-]
A number of years ago when I heard about this, I was pretty angry and made a private fork of git immediately in which I tried to scrub away the SHA-256 bullshit. But that's basically just paddling upstream with a spoon for a oar.
The stewards of Git are going to do whatever they want, and there is nothing you can do about it if you don't have the clout to create a fork that takes the lead.
No amount of discussion will do anything because they've already decided that their view of the situation is correct. Git hashes are not just content identification but a digital certificate mechanism, and their collision resistance is a grave issue that must be fixed, the end.
You will be browbeaten in any discussion; it's not worth the energy in a world replete with issues.
SmasherEpilepti 2 hours ago [-]
I find it disappointing how few people are addressing the proposed "Independent Tree Hash Headers" solution. It's probably the most interesting part of the article, but it's getting the least attention.
I came in expecting to disagree strongly with the article, but ended up agreeing more than I didn't (though I still don't 100% agree, as collisions are still an issue for mirrors). I find the concept of multiple hashes per commit quite interesting. It would allow mixing hashes in one repo, wouldn't break submodules, and tooling could be used to reject commits without any secure hashes for a gradual transition (like enforcing signed commits/tags).
juliusdavies 15 minutes ago [-]
Meh. Git 3 can still init with the old sha1 hash.
Also GitHub should auto-detect the format on the very first “git push” and handle it then. If they actually require config to be set appropriately during the initial “new repository” dialog that’s just bad ux.
As for this being a looming disaster for industry… it’s inconvenient. And lots of things will break. But we will survive and come out the other side I’m certain. We converted from the Julian calendar to the Gregorian calendar 500 years ago. Surely we can handle this.
kccqzy 3 hours ago [-]
I agree with the main thrust of the article, that we should instead trust the transport mechanism rather than the cryptographic properties of SHA1, but because of this I don’t really think switching the default will be a costly mistake. I rarely use full SHA1 hashes right now; I only use the truncated version and I don’t think any user cares about the length of the full hash. As for compatibility with forges, it’s just a small UX problem that should be solvable: don’t let the user choose the format when a repo is created; instead choose it when the first push happens.
sigmar 12 hours ago [-]
>it will be an incomprehensibly expensive and ultimately valueless and avoidable global nightmare.
thought "costly" in the title and "incomprehensibly expensive" in the subheader meant this piece would discuss how much less performant sha-256 is on modern machines, but didn't see anything. isn't there hardware acceleration? how much worse is it?
schacon 11 hours ago [-]
Actually, I think sha-256 is possibly faster than the sha1dc variant that Git currently uses.
I just sent a patch series to the list that enables sha1dc to be accelerated on modern CPU architectures to close to normal SHA1 speeds, but since it was ported from a Rust project by an agent, it will never be applied.
Last I checked SHA-256 was faster than SHA-1, and SHA-512 was even faster (though the output is annoyingly long).
mike_hearn 12 hours ago [-]
He means costly in terms of human effort and wasted time.
storyinmemo 12 hours ago [-]
Yes every repo is either one or the other but you fix that by rehashing the entire repo. Everyone can do this independently. It's entirely possible to maintain to identical repos in SHA1 and SHA256 mode but for the most part I suspect once updated people will simply pull down the new repo and use git 3.0 as a required version.
As migrations go, it's reading as simple to me. You'll just have to backpoint the commit signatures. I must assume there's a backwards compatible reference for them in git 3, right?
Or drop them and reference the old structure in a dire pinch.
mort96 12 hours ago [-]
Re-hash the entire repo as in rewriting all history? Hooo boy will that be a mess, I deal with things which reverence commits by hash in repos all the damn time. There are thousands of them in every Yocto project!
iamnothere 12 hours ago [-]
Yes, this would cause big issues for Nix based build systems or any others that reference commits by hash.
ba1afd89f34cb23 11 hours ago [-]
Do you have any external references to any commits that matter, for example in your communication platforms (emails, Slack) or your bug tracker? Or, worse yet, in places where they aren't just text format references, but used for things like CI/CD caching decisions or security scans?
Once you rehash the entire repo, every single one of those external references will be broken. Because no, there's no support for looking up old hash -> new hash or the reverse.
a1o 4 hours ago [-]
I think I remember something for this for mercurial to git migrations, I hope when its git 2 to 3 something similar is made (or it will take some time to adapt like when python did its 2 to 3 migration).
xd1936 12 hours ago [-]
"The migration to the new format is simple; Just re-write everything in the new format, but also keep the old format around forever too since data is lost in the new format!"
bmacho 10 hours ago [-]
I don't agree that SHA-1 is much longer feasible for git.
But I also don't think that switching to SHA-256 must be painful. A git2->git3 converted repo could just store all the past hashes, so existing links don't break.
Buttons840 5 hours ago [-]
I was thinking the same. Can't the git CLI see a hash and say "well, I don't see any matching SHA-256 hash, but let me check the Legacy SHA1 hashes I have stored", and still resolve an old SHA1 hash to the correct commit?
hedora 2 hours ago [-]
It could even do that + barf if it saw two objects with matching SHA-1 but mismatched SHA-256!
colinublake 1 hours ago [-]
the sha1→sha256 flip is gonna break every script that assumes 40-char hashes. my own deploy glue does that lol. not looking forward to the grep day
benthecarman 5 hours ago [-]
Core of the issue seems like github UX issues that they can solve
hedora 2 hours ago [-]
In fairness, there's a second issue: Git repos should just support multiple hashes, so you could just add SHA-256 to a SHA-1 repo, which would then transparently support both SHAs.
This even would make the SHA-1 git objects collision resistant when stored on a trusted server, even with untrusted clients. (Exercise left to the reader.)
rurban 2 hours ago [-]
I'll probably switch to git-evtag then, and keep the old SHA-1 then. Same as Google.
kazinator 11 hours ago [-]
I positively don't care about the collision issue.
If you need to certify the authenticity of some code, and you've decided that a Git hash of any kind is going to be your certificate, you have a problem between keyboard and chair which is not fixable by stronger hashes in Git.
I don't want instability and churn in tooling.
Strilanc 11 hours ago [-]
The post's argument that hash collisions are irrelevant in practice is not convincing at all. Basically they amount to:
1. Collisions aren't as bad as preimage attacks
2. Even if you made a file-with-malicious-hash, how would you get people to pull it?
3. Other attacks are a bigger problem (social engineering)
(2) is laughable in a world with github. It's common for unknown people to submit pull requests to code bases, and for those changes to be reviewed and merged. For example, as part of reviewing pull requests, I have `git fetch`'d proposed changes to my local machine to check behavior on some additional test cases. "If you fetch it you're fucked" is unacceptable as a security boundary.
(1) and (3) are just tu-quoque arguments about other attacks being worse. The relevant question isn't how bad other attacks are, it's how bad this attack is.
The fundamental problem with collisions is that software often assumes they can't happen (or is not tested against them). Thus collisions can trigger bugs, or otherwise cause surprising behavior. For example, webkit figured the colliding PDFs demonstrating a sha1 collision would be excellent for unit tests, so they merged the PDFs into their SVN repo... which completely fucked it [1]. I don't know the exact internals of git so I can't comment on how you would get surprising things to happen, but "oops the file you merged was different than the file you reviewed" and "oops the repository got corrupted" seem entirely plausible.
(2 counter) is impractical because all nodes of git will not replace objects if it thinks it already has it. So any attack has to assume this is the first time the node fetched, which is difficult before trust is established, which is difficult. This is part of the argument Linus originally outlined for this vector, which is that it only works for _very recent_ objects.
(1/3 counter) is not what I argued. I argued from the worst-case position that collision and preimages were theoretically cheap and fast. Even in that case, I feel my arguments hold.
The main issue here is that you assume you can replace an existing object with a replaced one, which you cannot. Not only that, but in all known cases, the sha1dc variant of SHA1 that Git uses will even _tell_ you that someone tried to do this, which singles out the source quickly.
6thbit 11 hours ago [-]
couldn't github reject a push that contains an existing hash in the repo?
schacon 11 hours ago [-]
It doesn't reject, but it will not replace. Same for a fetch/pull. That is another issue with this attack vector (that Linus also mentions) - it has to be the _first_ time that a node has seen this object. It makes the attack even more difficult than it already is (in like 4 different major ways)
pavon 12 hours ago [-]
Ugh, I didn't know that SHA-1 submodules wouldn't be supported in SHA-256 repos. That changes the transition from painless to a major dumpster fire. Having to maintain converted forks, and use different hashes from upstream is going to be a mess.
hnlmorg 12 hours ago [-]
git submodules has always been a major dumpster fire. You’re honestly better avoiding regardless of SHA-256 incompatibilities
OkayPhysicist 12 hours ago [-]
Can someone more cyber-pilled than me explain what the actual risk with Git hashes being susceptible to collision attacks is? Obviously accidental collisions are problematic, but to my understanding the probability of that is still approximately zero.
Best I can tell, all a forced collision would do is let someone who already has control of a repo modify the history in a far from plausibly deniable way. Which in practical terms, they already could do simply by replacing the whole thing, because who's out here using git hashes as a security tool? Every pinning I've ever seen has been to tags (which can be modified at will), or hashes of the actual payload (which doesn't need to be the same as what git uses).
edelbitter 2 hours ago [-]
> who's out here using git hashes as a security tool
Among others, dependency management in Rust [1] and Python [2] sometimes uses references that work similar to https://github.com/rust-lang/rust/commit/ec999ed [3] to suggest one particular version of the project, authored by the specified maintainer.
Unfortunately, it means neither, unless you pushed it. The hash points to whatever the first person uploading it to github submitted. And the author/org name in the URL is window dressing: all the objects go in one big bucket regardless of push permission to one particular fork (because why wouldn't they - today, collisions are believed to be recognizable because the cheapest way to craft them results in clear tells).
[3]: N.B. the "This commit does not belong to any branch on this repository, and may belong to a fork outside of the repository." warning Github has started to add to URLs like that.
Palomides 12 hours ago [-]
I see commit hashes used all the time, like in yocto recipes for example
r3trohack3r 12 hours ago [-]
> We can go through years of this SHA-1 to SHA-256 migration and then quantum computers break 256 and we're back in the same stupid boat again.
SHA-256 is considered quantum safe by the NIST and is left out of PQC migration guidance entirely.
schacon 11 hours ago [-]
It was theoretical - the point was that maybe some paper is published or some new tech or issue comes up. Now we have to do this again. If we separate the concerns, then we don't have to deal with both as though they're one problem. We can deal with one thing for content addressing and another for trust and security.
ghusto 11 hours ago [-]
Until it isn't.
Things like this have a tendency to to be revised as time passes.
TheRealPomax 11 hours ago [-]
For now. Turns out the pigeon hole principle still holds.
AndrewDucker 9 hours ago [-]
1x10^77 is a lot of pigeon holes.
TheRealPomax 7 hours ago [-]
Yep. So was 1.46 x 10^48
flowerthoughts 10 hours ago [-]
Oh, agreed this sounds like a terrible migration path and shouldn't really be needed in the first place.
What I'm missing in the article is whether any Git server accepts replacing a SHA-1 identified object it already has. If it doesn't, then the distribution trust discussed holds, and keeping SHA-1 seems fine. Adding additional signatures seems fine for those who need transitive trust.
12 hours ago [-]
kazinator 11 hours ago [-]
Make git init use SHA-256 if git is invoked as git3, SHA1 if invoked as git2.
Plain git init could fail with a diagnostic: informing to use one of the two aliases or an option.
What people don't want is making git repos SHA-256 by accident and finding out later that they made repos not compatible with older git.
metalliqaz 11 hours ago [-]
just make SHA1 the default in all cases unless the user specifies otherwise
kazinator 11 hours ago [-]
I agree, but you're not going to sell that argument to a herd which has decided that git hashes are digital certificates which must be replaced with SHA-256, or the sky will fall.
njt 12 hours ago [-]
schacon: Really like the "Independent Tree Hash Headers" idea.
How difficult would this be to get this functionality into git?
Would it cause any breaking changes with older versions?
Have you discussed this with any git devs to see if they are open to adding it?
schacon 11 hours ago [-]
Actually, this entire blog post came out of a short chat at Git Merge a few weeks ago with Jeff King. I argued more or less this and he didn't _entirely_ disagree, though he has good counterarguments on the list over the last few years, so I don't really know how he thinks about it ultimately.
I would write this to the mailing list, but I thought a conversation that includes people outside that list is more interesting to me. Ultimately I'm not sure if I'm dumb about this or the whistle blower that's willing to actually say "maybe this isn't the right call"
schacon 11 hours ago [-]
Also, functionally, this is incredibly easy to add to Git.
bawolff 4 hours ago [-]
I think the only good argument here is that sha is maybe not a security control for git. I think every other argument in this article is incorrect
a) it's relatively fast and impossible in a practical sense for two different files to accidentally hash to the same value.
That is silly. We are not worried about accidentally triggering. We are worried about intentional triggers.
I dont know why people always bring this up for hashing. In any other context it would be considered silly. If someone said, the chance of triggering a buffer overflow by accident is low, we would call that silly as we aren't worried about accidental triggers.
b) second pre-image vs collision.
In a world of open source where we accept commits from randoms on the internet, i think collisions are just as relavent as second pre-image.
eviks 3 hours ago [-]
a) it's silly to stop reading at that quote because the following text deals with "identification"
b) so, you agree with the blog?
"So, any realistic interesting attack vector therefore relies on a collision attack,"
bawolff 2 hours ago [-]
> b) so, you agree with the blog? "So, any realistic interesting attack vector therefore relies on a collision attack,"
My reading of the blog is that they are dismissive of collision attacks. In context of git, i disagree. I think there are plausible attack scenarios involving collisions, or at least, just as plausible as second pre-image.
If you mean do i agree with the blog that impossible attacks aren't possible? well yes obviously, but i think that goes without saying.
Graziano_M 11 hours ago [-]
I suspect everyone renaming their branch from master to main caused more unnecessary breakages and toil than this ever will.
jcranmer 11 hours ago [-]
The problem is existing repositories have SHA-1 commit hashes, and if you also end up changing those...
There's lots of tools that refer to git commit IDs. Some of those tools may even hardcode a commit ID to be 40 hex digits long. The fact that these tools are external also means that "oh, just rewrite the commit messages or code to refer to the new IDs" isn't feasible. The only way to not break the world is to let people refer to existing commits with their SHA-1 hashes in perpetuity, and it doesn't sound like git is set up to allow this in any way, which means that existing repositories have to stay SHA-1 in perpetuity and that will cause fun down the line if you start having to make SHA-1 and SHA-256 repositories.
Changing from master to main is a one-off change. It might require changing your scripts once to refer to 'origin/main' instead of 'origin/master', but other than that, there is essentially nothing more that needs to be done, there is no risk to historical artifacts that needs to be mitigated.
bkolobara 12 hours ago [-]
I run a small git/jj forge and for us it's already painful dealing with this. Can't imagine how GitHub is going to handle it.
JaumeGar 12 hours ago [-]
The xz backdoor is basically his point in practice — that was a maintainer-trust compromise, not a hash collision.
mdavid626 10 hours ago [-]
My prediction: 20 years from now everyone will still use SHA-1 git. That will be simply easier.
Levitating 10 hours ago [-]
Do you want to bet on that prediction
seebeen 7 hours ago [-]
I would
gfody 2 hours ago [-]
> not really practical to exploit in any demonstrated way
like gitc0ffee?
limonkufu 6 hours ago [-]
It seems people are missing the point: it's not even the submodule incompatibility that's going to become an issue majorly (like python2 --> python3 but worse), the main issue is the loss of traceability for repos that changes in place (which I assume many will do). Imagine what will happen to these:
- SLSA and Provenance or SBOM data in the supply chain security that uses commit hash. All the previous images are now pointing to a non-existing commit
- All the documentation and tooling as the article calls out
- All your traceability links from your project tool to your git repo, they will lose all the past data as it will be dead links
So I hope there IS NOT a migration path for in-place replacement!
nixpulvis 10 hours ago [-]
I'm going to completely ignore the first part of this post because I'm not interested in arguing about how severe the issues with SHA-1 are. I think it's accepted that there are flaws.
So given that, I'm more interested in the arguments for why migrating to SHA-256 is problematic.
The biggest issue I see, after skimming over it, is the submodule breakage for new projects trying to link to old projects. This seems solvable frankly, but is the only serious issue I see. Everything else will be worked out as software is updated IMO.
wat10000 3 hours ago [-]
I feel like if it takes this much ink to explain why using an insecure primitive is actually safe, you should just fix it.
The arguments make sense, but how ironclad are they? How confident are you that some clever black hat won’t figure out a way to take advantage of it?
This is one of the most widely used programs in the world. Let’s close the hole.
Magicrafter13 12 hours ago [-]
The first reason the author lists for why this will be bad is only an "issue" on Git hosts that don't allow repo creation on push (which is brain dead of GitHub). Any other host, you push your new repo, and it will see the hashing algorithm, and receive the contents accordingly.
Submodules is a legitimate argument against this, though I don't know how widely this feature is actually used, and similar to the arguments in favor of switching the default branch from master to main, this is simply a setting which can be changed.
I do like the idea of commits having both hashes, and am surprised that idea has not been explored further.
Generally though, I think the author's strongest argument is simply that the change isn't strictly "needed", and all the other issues presented aren't the strongest arguments against change.
kittikitti 2 hours ago [-]
So many people coping.
thunderfork 12 hours ago [-]
A lot of replies here seem to be asserting that this "isn't that hard" without addressing the thing that makes it most hard: submodule compatibility and the breadth of tooling
pixl97 12 hours ago [-]
Submodules are a mistake.
mort96 12 hours ago [-]
They're the best way we have to reference other repositories from one repository. All other solutions don't have the benefit of being built in to git and having support built in to all git forges.
Ecosystems like Yocto are built around having meta layers as submodules. And, despite the usability flaws of submodules, it works really well.
I also use submodules to include dependencies into C++ projects a lot. It works fine.
bryanlarsen 12 hours ago [-]
git subtree and git subrepo are compatible with all git forges and don't require normal developers to install the extensions. Only the person/bot doing the occasional sync to the external repo has to install the extension. I prefer git subrepo for most (but not all) use cases.
pavon 12 hours ago [-]
Note that subtree and subrepo have the same SHA-1/SHA-256 incompatibility issue that submodules do, so this will be just as much of a trainwreck for them as well.
mort96 12 hours ago [-]
What's the advantage to using git subtree or git subrepo instead of git submodules? I've never heard of this, what's the difference between them? If it's an extension, how do people without the extensions end up downloading the code from the other repos?
How does it work with MRs, can I submit an MR which consists of changing the referenced SHA (and have it not show up as changes to every file in the referenced repo)?
bryanlarsen 12 hours ago [-]
They work by copying one repo inside another and providing tools to copy/sync it back out again. It's not a link, it's a copy. It's almost the same as copying the files into your repo and git add'ing them, but there are accounting and tools to pull changes from the subrepo back to the external repo.
The trade-offs are relatively obvious. It'd be a poor option for Yocto, but is a better option for most corporate repos.
mort96 12 hours ago [-]
Oh, I didn't want to vendor another repo into mine, I just want to store a reference to it. I'll keep using submodules then, as they're easier to work with than tools like gclient and repo.
I really don't get the hate. They're not hard to work with. Just a bit shitty UX but if you're using Git you're used to that already.
ghusto 11 hours ago [-]
I won't defend submodules, but I also don't accept this as a response because it's irrelevant. They are used and it will be an unbearable pain when they break.
purpleidea 12 hours ago [-]
> Submodules are a mistake.
Someone started this FUD a long time ago and it has worked. Instead of using an elegant mechanism, project have built inelegant wrappers on top of git like go.mod which are actual mistakes.
okanat 6 hours ago [-]
It is the same people who say Git is unseasonably hard to learn. It isn't. Git and its command line is ugly and hard to remember. Understanding what those commands do at a commit and a branch level isn't really difficult. Even the internal object model is only moderately difficult.
I say this as a person who strongly dislikes many many aspects of Unix and Linux due to bad design and terrible UX. Git has a better design than any Unix program you get.
Submodules work okay. It is just Git LFS but for Git repos. Get over it.
hnlmorg 12 hours ago [-]
I’m not the author but I agree with their opinion.
Compare the UX of go mod with git submodules. One is easy and the other is about as fun as having teeth extracted.
git’s UX has never been its strong point. But submodules takes that pain to a whole new level.
pphysch 12 hours ago [-]
I find them extremely useful. It gives me a monorepo experience in repos that otherwise aren't/can't exist as one monorepo for various reasons.
theowaway 11 hours ago [-]
they could just have taken the sha1 of the sha256.
GrantMoyer 5 hours ago [-]
Then `git fsck` wouldn't be able to tell if an object uses the sha1 scheme or the sha1∘sha256 scheme, so it may need to compute the both hashes to check an object's integrity. Also, an object name alone wouldn't indicate that the weak scheme shouldn't be used to check integrity, so a malicious sha1 object could be swapped in in place of an sha1∘sha256 object (if second-preimage is found).
jmyeet 6 hours ago [-]
I'm honestly still shocked any of this happened.
Prior to SHA1 we had MD5, a decade earlier. MD5 collision attacks had already been widely documented and known. It was the most obvious thing on Earth that this would happen to SHA1 too. Apparently, Linus never realized there was a need for cryptographic security and that the hash was purely internal.
Here's what I honestly think was a factor. I think C programmers fell in love with the implementation that you could throw around a fixed hash record on the stack. It's incredibly efficient. But it's an efficiency that doesn't really matter because as soon as you read from or write to a disk or a network or even memory, any cost saving is completely gone.
More than a decade ago, some people wrote a Java implementation of git (jgit?) and despite all their optimizations, it was (IIRC) only half as fast as C git. It is of course because Java at the time had no concept of stack values for non-primitive types so couldn't compete. Personally, I was impressed: only half the speed? That's pretty good.
For something that's only 20 years old, the Git SHA1 assumption is some of the worst technical debt we have in the modern era.
Here's another thought: when people make a lot of these programs, they often make the mistake of not separating the program version and the network protocol (or just the external API). So you end up with brittle client-server implementations where you have to upgrade both the client and the server at the same time because they lack a network abstraction.
The other end of the spectrum is video streaming where you have codex, container formats, transport protocols and so on.
What a mess.
PunchyHamster 6 hours ago [-]
> I pull it from there because I trust that GitHub has its authentication game together enough that it's unlikely that anyone malicious pushed something there without the maintainer's knowledge.
Hahahahaha, that's some level of delusion
ltbarcly3 12 hours ago [-]
This seems like Y2K fud.
The alternative to making sha256 the default is to leave sha1 the default. Nobody changes to sha256. sha1 is broken in 10 years. Suddenly everyone has to switch all at once on the same day because it is a critical security issue, but github never implemented sha256 because they didn't have to. This would be a major problem.
This is very very easy to fix if you run into it.
1. Adopt git 3.0 if you can with sha256.
2. If you can't use sha256, set the config to put things back to sha1. Wherever you need to do this you probably already set dozens of ENV vars or settings, just add a new one.
Or write a 15 page analysis about how the above is so hard people will probably just find it catastrophic to even think about.
wavemode 12 hours ago [-]
> sha1 is broken in 10 years. Suddenly everyone has to switch all at once on the same day because it is a critical security issue
If you read the OP article, the entire point he's making is that this would never happen, because a hash algorithm being "broken" doesn't matter in practice, because true supply chain security has nothing to do with file hashes.
bityard 12 hours ago [-]
It's easy for _one person_ to fix. It's not easy for the entire git ecosystem as a whole. GitHub, large internal corporate git repos, CI/CD systems, projects with submodules, etc. The second half the article explains all of this.
iamnothere 12 hours ago [-]
It’s already broken, but even though it’s broken it’s hard to generate git collisions because of the repo metadata. It’s easy to generate (for instance) standalone PDFs with identical hashes, but doing this with git in a useful way is much harder.
That said, it’s still a good idea to migrate to a more robust hashing algorithm. Defense in depth, etc. Just because it’s a difficult migration doesn’t mean it shouldn’t be done.
schacon 12 hours ago [-]
You can certainly do this, as I said, this is Google's backup plan. But defaults matter. People will start running this and getting repos that are uselessly incompatible with other repos, tools, libraries and server instances. Having it as an option is one thing. Making it a default will cause a lot of pain for people who don't want to care about this.
ltbarcly3 12 hours ago [-]
"defaults matter" is an argument for this change, not against it.
schacon 12 hours ago [-]
No, my argument is that the change should not happen at all and nobody wants it and it gains the community very, very little but the default change is forcing it on everyone and most will be _entirely_ unaware - now having to solve problems that are difficult to understand. Defaults also matter when they are the wrong defaults.
ltbarcly3 12 hours ago [-]
This is not difficult to understand. It's very easy to understand.
addaon 12 hours ago [-]
> Adopt git 3.0 if you can with sha256.
Who is "you" in the context of a distributed version control system? I think this is not just the plural you, but the unbounded you -- it's all people who not just interact with your project now, but who you hope may interact with it in the future. The question is what the cost is of committing a near-infinite population to this migration, not the cost of doing a single `brew update` on your personal machine, no?
ltbarcly3 12 hours ago [-]
Just clone the repo again. Jesus Christ, you act like the simplest thing in the world is some kind of insurmountable challenge.
OutOfHere 12 hours ago [-]
For the record, Y2K was not fud. It was very real, in a long list of datetime problems that are to come. Further datetime problems are coming at scheduled dates.
bigstrat2003 12 hours ago [-]
It was definitely FUD. There was a real problem (date counters would roll over), but the impacts of it were so ridiculously overstated that it eclipsed any sane discussion of the issue. We had people at the time predicting that planes would literally fall out of the sky when Y2k hit, which was never a realistic possibility.
pixelesque 12 hours ago [-]
> which was never a realistic possibility.
Because a lot of work was done to prepare and fix potential issues.
OutOfHere 10 hours ago [-]
That's the problem with deniers. When responsible persons take preemptive action to prevent tragedy, like with Y2K, the diners say it was FUD. When people don't take action, like with climate change, they say it wasn't important considering it's not them who's dead, totally discounting those who have suffered or died as a consequence. In summary, the deniers are so incompetent that they can't be trusted to correctly maintain a car, let alone civilization, considering they would never even the replace the necessary parts at the right schedules in their car.
globular-toast 11 hours ago [-]
Thanks for writing this. I'd only been loosely following it and I hadn't realised how bad this is going to be. I have repos with tens of submodules and it's going to be a nightmare if any of them switch to sha256 in place. Not to mention I won't be able to use any new projects unless I rebuild my repo and all the submodules therein.
I thought the master to main thing was bad enough but this is going to suck. And just like the master rename it achieves basically nothing.
What is it about these projects that attracts people who just want to change things for the sake of it? Real engineering means coming up with a solution for backwards compatibility. This is just irresponsible and, frankly, a fuck you to everyone who will be affected by this.
Lumich 9 hours ago [-]
« people who just want to change things for the sake of it » — Not for the sake of it, but to “make the world a better place.” The intention is noble (well, mostly, at least let's assume it is). Of course, there is disregard for history and her deplorables (or not so deplorables), comes with being “progressive”. Which is why I like Windows better (not 11 though, will probably have to go back to Linux at some point).
« What is it about these projects » — Maybe that they're “at the forefront.”
seebeen 12 hours ago [-]
[dead]
mrtesthah 12 hours ago [-]
…
ande-mnoc 12 hours ago [-]
What does OpenAI have anything to do with this?
12 hours ago [-]
quotemstr 12 hours ago [-]
Would the author feel the same if git had used MD5 instead of SHA-1?
schacon 12 hours ago [-]
I do actually literally write in this that if it was MD5 it also would not be a problem.
quotemstr 11 hours ago [-]
Fair cop.
happytoexplain 12 hours ago [-]
They address this very theoretical. In short: Yes. Which makes sense if you don't treat the hash as a form of security against malice, especially in the case of attacks that are already impractical, which is the entire thrust of the article.
eviks 3 hours ago [-]
Follow the ethos of the quote master!
> We could be using MD5 and it would honestly probably be just fine.
As linked by another commenter in this thread, Linus worked out years ago that even if someone inserted a malicious object into the kernel repo, it would at best be a nuisance and not a major concern.
12 hours ago [-]
pasteleft 2 hours ago [-]
Didn't GitHub broke the entire CI system by switching to "main" branch :thinking_face:
Anyway, I think it'll be the same.
Tools will support SHA-256 quickly and we might have a migration program that converts SHA-1 repo to SHA-256 repo.
The only problem is that git (and related tools) will get twice as big...
crispr245 2 hours ago [-]
Claude, make the hash use SHA-256 rather than SHA-1. No errors plsss.
Besides the possible implementation/deployment issues they will or will not face with this update, I can empathize with the idea that of not wanting to have a possible vector of attack in your system. Particularly today with AI being able to find novel exploits, I could see a future where a vulnerable hashing system leads to a malicious injection attack.
The author argues that "If I wanted to get untrusted code into Android, it's so much simpler to bribe or convince the maintainer of a popular downstream project" which is a really a red herring in this matter since that is literally a completely different issue that obviously no software update could ever fix.
Nonetheless I do agree with him in regards of how complicated and messy this whole process will be. Crypto migrations have been historically difficult, expensive and overall ugly, but not impossible...
1) It's claiming SHA1 insecurity is theoretical, while SHAttered from 2017 was specifically a pratical proof of concept. The only reason Git wasn't affected, is because they didn't bother bruteforcing a git-blob prefix.
2) It's claiming collision attacks don't matter, only second-preimage attacks do. This is incorrect, collision attacks are enough for code-smuggling problems, when two repositories are on the same git commit (verified by the full commit hash), yet contain different code in their git checkout.
3) The Linus quote "The real security is in distribution" is arguing that "git's content-addressed system should not be used to address content". It's arguing that, in case of curl|sh, you shouldn't use a sha256sum-gate to pin the content to something you've reviewed, you should instead ensure curl is fetching from an https server.
2) I specifically argue that even if both attacks were practical and cheap, it's still not the problem we should be focusing on.
3) Have you read this email (that I linked to)? It is almost the same general message (20 years ago) that this blog post is. It literally goes though a theoretical object replacement attack and how dumb this scenario is and so SHA-1 is fine.
https://lore.kernel.org/git/Pine.LNX.4.58.0504291221250.1890...
It seems unlikely it will stay that way forever. Typically attacks get more efficient over time as researchers find improvements, not to mention computers getting better.
In 2015 it was estimated to cost $100,000, now the estimate is down to $10,000. Where will it be in 2035?
Then you would have security researchers making conflicting claims depending on which repository they first pulled from, even though they are on the same git commit hash.
... for Linux
... and developers working for it constantly
the attack wouldn't work. Joe Schmoe? It's worse than just "being compromised"
You have repo of dependency locally, let's assume you downloaded good copy, the commits get compromised, you're safe.... right ?
Nope, if there is build server along the way and ESPECIALLY if it practices building from clean state every time, the build might be infected while your local copy is clean, giving no chance to notice it, unless your entire chain including local builds are reproductible AND you actually check it
Nothing else matters.
Git hashes are not supposed to be a security mechanism. If your basis for trusting that you have the right checkout is the git hash, in a situation where you have legitimate concern about untrusted parties manipulating remote repositories, then you're simply wrong.
When I check out code from a git repository in a pipeline using a git hash, I expect the code to be exactly what has been reviewed by me under that hash.
Everything else would just be a crazy invitation to make supply chain attacks uncircumventable.
Well, what about someone who is fetching the commit from that server for the first time and has nothing to compare the hash against?
Oh, that would never be a problem for widely disseminated, popular, open source project, so it doesn't matter.
Or if first writer wins, and I know that you have a popular non-GitHub repo that you're about to migrate into it, then I could pre-poison the namespace by writing my own version of a commit that I see you already have in Codeberg or Savannah or wherever.
I don't swear that this is how GitHub actually works, but I've had knowledgeable friends swear up and down that it is. And honestly, it'd make sense. They could shard storage by the first 4 digits of the hash or something, and that'd be vastly more efficient if all commits were writing to the same space.
Turns out securing a service to transfer a single hash is a lot easier than securing a service to transfer gigabytes of data.
Even if I don't fully trust Github, it is still incredibly convenient to be able to upload my code there and then send someone an email telling them to fetch commit `123abc` from some repo link. As long as my email isn't compromised, that should be secure.
Note that the US CLOUD Act means that, if someone figures out how to actually use collisions to compromise that CI machine, then, if the US government asks Microsoft to do use that vector to break into an overseas machine, then Microsoft will be legally obligated to do it.
But you still need SHA256 for that
Probably a naive question, but why not kill two birds with one stone if it can be done for a reasonable cost?
A SHA-256 sum, though very good, only assures you with great confidence that you're looking at the same thing you looked at before, or that someone else is looking at elsewhere.
It is not a digital signature, and we don't want digital signatures to serve the role of content hashes.
Speaking of signatures, we have support for them in Git; you can use gpg to sign commits, and set it up to be done automatically.
Nobody is going to fake your commit such that the fake has the same SH-1 hash and your GPG signature.
The worry there is that the key holder (whether the legitimate one, or a malicious party who got a hold of the key) somehow does this: creates a new commit, signed with their key, which somehow has the same SH-1 as an existing signed commit. The git hash includes the GPG signature, so there is a significant layer of difficulty there which is likely harder than faking an unsigned SHA-256 commit.
The attack is I pre-author `Makefile => foo: echo "hello"; bar: echo "world"` along with `Makefile => foo: echo "hello"; bar: rm -rf / ; /* $ELDRITCH_SHA1_SPIRITS_GO_HERE */` that both hash to `ff1234...`
I then prepopulate the repo with `echo "hello"`, wait 6-9 months, then submit a commit for `echo "hello" ; echo "world"` and keep (in my back pocket) the alternate implementation that also includes $ELDRITCH_SPIRITS to force a collision and MY predetermined change in functionality.
I then have free choice as to whether I serve them "hello world" or "hello && rm -rf", and THAT's the plausible problem to avoid: the ability to "cloak" content anywhere within the repo if you have enough $ELDRITCH_SPIRITS and GPU's.
You have _really_ good points, but are woefully confused. The proper answer is (would have been) to include `tree ff12354...` along with `tree-sha256 abc123456789...` for another 20 years along with a `[git.hash_strictness]: default/lazy/strict`, and some oddball `git-rerere` type packfile extension which lets you map `sha1:ff1234... => sha256:abc123456789...` "transparently" rather than the horrific situation you're laying out (correctly!) that forks the ecosystem in to "longhash" and "shorthash" when most repos don't even care in the end.
Yes, I didn't understand that the GPG signing just operates on the top level object in the commit and trusts the SHA-1 hashes contained in it.
The signing process doesn't recursively traverse the bytes of the commit to pull them into GPG, like you would expect.
It's like, imagine you made a "bill of materials" of your project's files consisting of their names and CRC-32 checksums, and then signed this file, and called your project securely signed, LOL.
This aspect can be fixed without foisting new hashing scheme into the content tracker. In fact, it must be fixed; users on SHA-1-based repos deserve secure signing.
It's really sneaky that the SHA-1 business (not intended to be a security mechanism) was embroiled into the signing implementation; that GPG is demoted to the strength of SHA-1.
Was that just to save some cycles? It's certainly faster just to sign the commit object!
We can round up the bits that make up a commit in a SHA-1-based repo, and sign those bits securely; this is a thing that is possible.
Yes but as my other comment explains, this is not particularly useful in and of itself.
You refer to PGP signed Git objects, but you also argue:
> Git hashes are not supposed to be a security mechanism
Guess what the Git PGP signature is signing.
The GPG signature signs some kind of hash calculated over the commit, minus the GPG header, which is thereby added.
The git hash is then calculated over the whole thing. The git hash is on the outside, and not part of the signing.
It kind of is - it’s signing the hash of the tree object, which is the actual thing that you’d attack with a hash collision
The actual git ‘tree’ object, which is the thing a commit actually points to, referenced by a hash in the commit. That is signed by the GPG signature.
That digest can be the SHA-256; since the infrastructure is there for it, signing should use SHA-256 regardless of what hash is used by the repository for identifying and linking content.
The "bytes passed to GPG" of course get hashed by GPG, using something better than SHA-1.
All bytes that comprise the commit should be hashed by GPG, rather than depending on the content referencing hash in the object tracking system.
This is something that is possible; it is not a logically deductive necessity that we just scan the topmost object and trust the hashes it contains.
It kind of is. Otherwise the whole idea of signing a commit with a backing git history (rather than just a snapshot of a working directory) collapses. The only guarantee you have that the git history is what is claimed by the cryptographic signature is some sort of Merkle tree structure. Either the original one, or you have to construct a whole new parallel one with a better hash, in which case, as I bring up in a cousin comment, why not just use a better hash in your original one?
Parent commits can have their own signatures, and it's true that there is an attack possible there where the same parent hash could point to two different commits, that have valid signatures of some kind (possibly from the same sneaky developer who is a bona fide project member).
Be that as it may, it's perfectly okay for the signature on a banana to validate only the banana, and not the gorilla that is holding it, or the vine the gorilla is swinging on, and the whole jungle.
It would be worth it to have better quality commit signing for SHA-1-based repos.
This is a significant degradation of the implicit guarantees given by a cryptographic signature, to the point that basically all personal use cases I have for signed commits would be invalidated.
Keep in mind that git does not have diffs as first-class objects. Every commit is just a snapshot of some state of the working directory. That means that without attesting to the integrity of the parents of a commit, the only thing a commit X signed by a person A says is "at some point on A's computer, the state of the repo looked like X".
Almost all the relevant questions I would want to ask are not answered by this. E.g. there is a malicious function F that is present in X. Did A write it? Don't know. Did someone else write it? Can't know for certain. Who introduced a certain feature? Don't know. Did A sign off on a new bugfix? Don't know.
All you know is that at some point the codebase looked like X on A's computer. A might not have made any relevant changes at all!
You can only back out a diff and therefore actually attribute a change to someone (either explicitly through `git blame` or informally by looking at git logs) if you have attestation of the parents.
The Merkle tree structure of git repos is interwoven through basically ever useful thing git does. Without cryptographic signatures implicitly carrying a promise of validity for that structure, this would make commit signing useless (depending on just how broken SHA-1 is) for the needs of any org I've ever worked at.
I don't think it's massively stupid. Unless you want to re-hash the entire Merkle tree structure to sign your commit, you basically have to trust the hashes in the Merkle tree (or have a separate parallel Merkle tree) at some point in what you sign, which means you do have to trust the SHA-1 hashes. Otherwise even with a cryptographic signature you can always spoof at least the git repo history (e.g. even if you try to directly hash the entire contents of the current commit).
Re-hashing the entire Merkle tree structure seems prohibitively expensive to generate (even with a lot of caching) and pretty complicated for e.g. verifying a signature. Or you can do that incrementally, but then you're just generating a whole new parallel Merkle tree structure.
Regardless, at the end of the day, you need to trust the integrity of the Merkle tree structure. And you can either do that by trusting the hashes of the current Merkle tree, or you have to completely recreate a new one with more trustworthy hashes, in which case why not just use better hashes in your original tree?
We use SHA-1 in our storage mechanism, which predates git. There is a (quite long) written threat model. Someone being able to generate collisions with an enormous amount of effort under just the right conditions is not a threat under that model.
Commit signing indicates otherwise.
The submodule problem is real, but the forge problem is entirely just "if forges implement support in a way that makes it painful it will be painful", and that's true of literally every feature that requires forge support.
The submodule issue is a different beast. But that one is currently a problem as well when a person choose a http url or ssh url. I have some extra git configs to normalize everything to ssh for instance.
"Both Fossil and Git started out using only SHA1 hashes. But when the SHAttered attack against SHA1 was published on 2017-02-23, the need to migrate to a stronger hash algorithm was recognized. Fossil added the ability to use SHA3-256 as an alternative on 2017-03-01 (six days after the SHAttered attack was first published). SHA3-256 is now the default for all new repositories and check-ins in Fossil, though older check-ins that occurred prior to SHAttered can still use their original SHA1 hash. Hence, no repositories had to be rebuilt and no hyperlinks were broken."
https://fossil-scm.org/home/doc/trunk/www/hundredandone.md
To me it's so interesting watching in realtime Git is still battling with this decision and for Fossil it was just another week of development.
That whole page is fun to read. Another fun fact somewhere else in the docs is that Fossil uses a grow-only set to store commits. They came up with this scheme some years before it was formalized by CRDTs!
Fossil isn't difficult to change not because it's technically harder for Git but because Git has a community and ecosystem that Fossil does not. The cost is not in the individual project for Git, the cost is because there is _so much_ in Git and this bifurcates everything.
To me it's more of a reality of creating a tool with a huge active community and a community of contributors and creating a tool with a small team and small community.
Is there any writeup on why it was easy for them and not for git?
From the skim I read of this article it seems both projects arrived at the same solution: support both but make SHA-256 the default.
:%s/git 3\.0/python 3.0/g
:x
Neovim's lazyvim plugin sucks because it takes over H C and L.
This is what I saw in one of the mails. > > There are organizations where SHA-1 is blanket banned across the board - regardless of its use
And also on git 3.0 breaking changes. > > SHA-1 ... recommended against in FIPS 140-2 and similar certifications
Since SHA-1 isn't used for security in git, they should've instead moved to a non-cryptographic hash function such as MurmurHash3 and avoid all these problems, instead of moving to SHA-256 until SHA-256 is broken and need to move to the next cryptographic hash that is now incompatible with previous versions of git repositories.
Linus' old argument was that the substitution would probably be noticed eventually, but that's specific to the way Linux uses git, and what he said probably isn't true in practice -- even if it is, there have been enough supply chain attacks since then to prove that even temporarily serving the wrong stuff to developers or CI is enough to allow lateral movement into other packages, production machines, etc, etc..
LWN had a good write up on this a while back: https://lwn.net/Articles/715716/
This is very likely the case. And if it is, then it's a lost battle. You simply can't reason with that kind of corporate people, let alone have an argument around this level of complexity. Kafka (the writer, not the message broker) predicted this 100 years ago.
When going through the article, my instinct was changing from "annoying" to "this really sounds like a Python 2/3 moment for Git" to finally "oof this is going to be a mess" in the libraries/submodules part.
> but the point is the SHA-1, as far as Git is concerned, isn't even a security feature. It's purely a consistency check. The security parts are elsewhere, so a lot of people assume that since Git uses SHA-1 and SHA-1 is used for cryptographically secure stuff, they think that, Okay, it's a huge security feature. It has nothing at all to do with security, it's just the best hash you can get. ... [1]
[1] https://www.youtube.com/watch?v=4XpnKHJAok8&t=56m20s
So Torvalds used SHA-1 purely because he needed a hash function with no other property than identifying content.
Sorry if the video answers this, but how does commit signing work if it doesn't rely on the hash algorithm being resistant to at least second-preimage attacks?
This flexibility ("agility") in cryptographic protocols is often seen as a mistake today, actually.
Which is sort of the hash algorithm approach git is taking with incompatible versions and a version break.
See also perhaps Wireguard, which touts itself as not having "cryptographic agility" because they wanted to avoid all (perceived) problems and complications of IPsec. But now that PQC is (allegedly) approaching there's no easy to update things because (AIUI) there's no negotiation possible in the protocol; you're basically standing up a 'Wireguard 2.0' that runs separately than the original.
> If an additional layer of symmetric-key crypto is required (for, say, post-quantum resistance), WireGuard also supports an optional pre-shared key that is mixed into the public key cryptography.
(From https://www.wireguard.com/protocol/.)
2. The hash function is not used for cryptographic purposes!
Was there “something like murmur” in 2005 that’s cryptographically better than SHA1?
Notably, a few of the featured author's reservations appear to be addressed. According to the Git docs:
- Objects can be referred to by their old, SHA-1 name or their new, SHA-256 name. This means old refs in docs and comments and such remain valid. The mapping between SHA-1 representations and SHA-256 representations appears to be intentionally bijective a.k.a. 1-to-1 (assuming no hash collisions), so that it could be re-computed on demand. The constraint of bijectivity appears to be the source of some limitations, ex. no mixed repos and submodules needing to match hash algroithm, but also bijectivity has strong benefits like the following items.
- A bi-directional dictionary is maitained from SHA-1 to SHA-256 names so translations between the two don't required re-hashing objects. This table could be recomputed on demand due to the bijection between names; it's only a performance optimization.
- A local SHA-256 converted repo (including an SHA-256 converted submodule) can interoperate with an SHA-1 only remote transparently to the remote server by translating names using the lookup table.
- SHA-1 based GPG signatures will be preserved. A commit can be signed based on its SHA-1 representation, its SHA-256 representation, both, or neither. The bijection means the two types of signatures are in a sense interchangeable, or in other words the bijection between object representations implies an equivalence relation on signatures. An SHA-256 converted repo can quickly validate an SHA-1 based gpg signature using the lookup table.
* we pin to commit hashes and expect this to refer to immutable content,
* commit and tag signatures are over the hash.
The Linux quote about trusting the distribution doesn't make sense to me, as Git is content-addressable and decentralised, although possibly at the time it was a reasonable position to take for kernel development, but [1] could equivalently happen for Git and the hash algorithm being non-broken is required for it to be noticed. Not having to trust the forge is a very desirable property.
[1]: https://lwn.net/Articles/57135/
SHA1-hashed objects should be able to refer to SHA-256-hashed objects, although this seems somewhat pointless.
But SHA-256-hashed objects should also be able to refer to SHA1-hashed objects, with a major caveat: if those objects themselves are part of a collision pair, then there is a genuine problem. But this is avoidable! Suppose that Linux decided to migrate to SHA-256. The upstream project could choose a pair of dates, say January 1 2027 and March 1 2027. Up to the first date, maintainers would be welcome to submit hashes of objects that are not yet in the repo but that they think they might submit later on, and, on that date, the upstream tree would finalize the list of these objects and reference it in the repo (with a new mechanism for this purpose). Effective the second date, the repo would start publishing SHA-256 commits and would never again accept a SHA1-hashed object that was not in the repo at the cutoff date or referenced as part of the Jan 1 block.
And now it would be impossible to get a new SHA1 collision in to the repo.
The only new git features needed would be:
a) actual compatibility so that a SHA-256-hashed object could reference a SHA1-hashed object
b) a new object type that's a list of allowed SHA1 hashes (or probably a tree of them) that is itself hashed with SHA-256 and a mechanism to link to one of these from a commit
c) a policy mechanism to set a repo to only allow SHA1-hashed-objects that a reachable from a preconfigured SHA-256-hashed commit
In any case, even if Git 3.0 were completely incompatible, it would suck, but it's not the end of the world. You just treat it as if you were migrating from one SCM system to another. CVS -> SVN -> Perforce -> Git -> Git 3.0 -> [...] been-there-done-that. This is something that both open-source and commercial projects have had to deal with over the years.
Or maybe it would be a repeat of Python 2.x -> 3.x. ¯\_(ツ)_/¯ With AI assistance, hopefully porting the tooling over may go a lot quicker and smoother.
[0] https://www.youtube.com/watch?v=eJJp0RE7cd4&t=1134s
The interop discussed is using copybara as a copy tool to move data from SHA1 based repos to SHA256 based repos and vice versa.
This is simply wrong IMO, for two reasons:
1. The attack on SHA-1 is a collision attack. Once you have frozen a hash, you cannot attack it with existing cryptanalysis. If there were a preimage attack it would be a different story.
2. Even if there were preimage attacks, one could freeze a mapping from SHA1 hash to SHA-256 hash.
In fact, #2 seems like en excellent design. Objects could reference such a mapping, and a repo could disallow conflicting mappings (the mappings would only be accepted if the mapped objects are reachable from the mapping and the mapping is correct).
I don't know what the solution is, but I'm inclined to believe that any repo with a single SHA1 commit is as weak as a repo with all SHA1 commits.
* I host a mirror of the Linux git repo.
* You download Linux from my mirror.
* You check out a commit, say fd179f8a05be3ccae366b9b96e176b51fbe54aab, which you know is a genuine commit through some out-of-band mechanism (mailing list, GitHub web interface, a line in a Nix file, whatever).
* You check whether the repository I gave you is legitimate or not by re-computing the hash of the commit which I claimed was fd179f8a05be3ccae366b9b96e176b51fbe54aab. If it comes out to be fd179f8a05be3ccae366b9b96e176b51fbe54aab, you know it's legitimate. If it doesn't, you know it's fake.
This is a completely normal use of Git. People download from mirrors all the time. People rely on commit hashes to identify a specific source tree. People trust that if whatever the mirror gave them hashes to the right value, it's genuine. That way, you don't have to trust the mirror.
If I can forge my own commits to have any hash I want, this whole model breaks down. I can replace some old commit in the repo with my own forged commit with the same hash, and when you download a copy of the Linux repo from my mirror, you'll receive a repo with malicious content, but it'll hash to the same fd179f8a05be3ccae366b9b96e176b51fbe54aab hash as a genuine repo would. This breaks the security model of Git.
That's a 160 bit hash, which is SHA-1, which has the security properties of SHA-1.
Suppose you check out a commit with a given SHA-256 hash. That commit object represent the root of a tree where all the edges are hashes (and types, etc). I'm suggesting one of two designs:
a) (Simpler but weaker) If Linus has published that commit, then he is confident that he hasn't pulled in any too-new SHA-1 hashes and that there are no collisions present in what he thinks the tree is. So, by induction on the traversal depth, there is only one actual object identified by each edge, and those objects contain the hashes of their child edges, so those hashes are all correct.
This breaks if there is a malicious collision already in the tree.
b) (Stronger but higher overhead and more complex) There would be an object or objects, discoverable from the root by following only SHA-256 edges, that encode a duplicate-free mapping from SHA-1 hash to SHA-256 hash. The client finds and parses that and then, as it traverses the tree, each time it reads a SHA-1 hash, it computes the SHA-1 and SHA-256 hash of the referenced object, verifies that the pair is in the mapping and also verifies that the SHA-1 hash matches what the edge requires.
I think that (b) is genuinely cryptographically secure in the sense that, if you can construct a commit that has the same SHA-256 hash as an official upstream commit but different contents, then there is necessarily a SHA-256 collision.
For B), I would think this could work, but it's a completely different solution from what you proposed and what I responded to.
How? Remember, there are (currently, anyway) no known SHA-1 preimage attacks.
> If I can forge commits with any SHA1 hash at will
We probably don't want to wait until there are practical pre-image attacks discovered to change away from SHA-1.
I think I stand by my second proposal. I also think it's absurd that, after all these years, upstream git still can't figure out a credible migration plan.
The "commit before" might be compromised, but the git commits refer a snapshot of a tree + a list of previous commit IDs, so the "new" SHA256 commit will not have any files altered
> But when you do want to push it to GitHub (or any host), you will need to tell them when creating the repository on the server that this is a sha256 project. Every project will now be in one bucket or the other and they cannot be mixed.
> This is immediately going to frustrate people because now they need to know what version of git they ran git init with and make sure when they go to GitHub to create the server repository, they choose the right one.
Github can fix this problem, and we have had siilar stuff in the past, when we had to switch repos. This is simple and doable, just need gradual work. I don't see the problem.
The git transition document talks about keeping a bidirectional mapping between old and new hashes, and converting between the two when pushing to legacy repos:
https://git-scm.com/docs/hash-function-transition/2.55.0
I suspect that’s what gitbutler.com means when they say that, although broken today, everything “will almost certainly be fixed” when git 3.0 ships.
While I’m sure they make a good argument in the hash theory section, I’m less inclined to believe the “train wreck” part of their post.
If you upgrade your house and leave the roof off then yes, it will be a “costly mistake” the next time it rains, but if your plans say you intend to put the roof back on then I’m not sure how I am supposed to interpret a blog post warning of the perils of a roofless house.
- It's implemented in a non-backwards-compatible way
- The benefits over the older model are a bit nebulous
- There's a large amount of tooling that needs to catch up, and little sign that there is movement there
(Yes us Hacker News users have plenty of use cases for IPv6, like self-hosting and peer-to-peer networking and so on; we are not the average user.)
This effect doesn't exist for the Git migration. Each repo can be updated independently; it doesn't affect users of other repositories, and most likely, the majority of devs will work on some SHA-1 repos and some SHA-256 repos with no issue.
If anything, I would compare it with the Python 2 to Python 3 migration, which was also painful, but succeeded eventually (despite being much less necessary in the first place).
Funnily enough, self hosting is why I can't use IPv6. I want vlan isolation, but only get a /64 from my ISP.
Fortunately the lack of IPv6 also isn't a meaningful loss anyway so whatever
The GitHub Actions ecosystem found out the hard way, through some rather high-profile compromises. They hotfixed it by adding "immutable tags" to their platform, and are now working on adding a lockfile to... easily reference a commit hash.
repo can also rewrite existing commit and you again won't be able to retrieve it so switching to commit IDs only lowers the level of failure somewhat
This is far from the case with IPv6!
Reality is that lots of ISPs do this right. And have for a very long time.
This is a good change, even though there is a massive technical debt in changing such a widespread system. It is worth the effort. Should generations from now still be using SHA-1 for their Git ops? Sometime you have to do the switch, otherwise you will never progress.
For my usecase basically i needed to know that every Git commit pointed at a cannoical blob. With SHA-1 you could generate two blobs which hash to the same SHA-1 hash, while you can do the format verification which helps i could not do that in my usecase as i did not know the underlying data. To fix this i had to very ugly have two methods of referencing any Git object, a cryptographically secure SHA-256 ID and the Git ID SHA-1.
This will be a train wreck. I hope they don't release before adding compatibility modes to keep the existing sha1's around in the database.
One example is to maintain git-replace refs for the rewritten SHAs but there also exist config flags to enable object format compatibility extensions that help translate the SHAs back and forth.
https://git-scm.com/docs/git-replace
[1]: https://github.com/newren/git-filter-repo
[2]: https://htmlpreview.github.io/?https://github.com/newren/git...
I don't think this is going to be a problem at all.
Not talking about archived code, but active projects that pre-date git 3.0
For massive perf and mem use damage. But oh well. And then we will wait for official version
But also collisions there aren't a big deal. People will cite short hashes when referring to things and that's not "broken".
If you're replacing the weakness of SHA-1 just by going to another algorithm, you better be prepared to go to the next one when sha256 collisions happen, and it doesn't sound like git's design would be easy to modify for this type of crypto agility.
I do like their proposal for using signatures to establish trust and allow swapping sha256 for whatever comes next.
However, it's not a git problem. It's an ecosystem problem. It's that every git repo has to choose one and they're entirely incompatible with each other. That is the cost and the difficulty.
It took 20 years to go from vulnerability in sha-1 to having to replace it out of caution. There is no such vuln in sha-256 yet. It could easily be 25 years before we find one, and another 25 years before we have to do something about it. Perhaps longer. Will git still be used 50 years from now?
All it takes is just one collision to consider it broken right?
But hey maybe the attempt to fix it makes git controversial enough it falls out of favor, and nobody uses it anymore in 2 years, problem solved? sure.
No, its considered broken before that stage. i.e. when someone discovers an attack that would allow someone to create a collision faster than they should while still being impractical.
> With the kind of compute power available nowadays and AI models I wouldn't be surprised we see it much sooner.
Computer power doesn't super matter, what matters is algorithmic breakthroughs. So far i dont think there are any examples of major breakthroughs of that type via AI, although perhaps i am just misinformed. Its still early in the AI revolution, it might still happen, but as it stands i don't think there is any reason to worry about that.
So this is the right time to post that everything they’re doing is wrong? Did you engage in all the discussions about it and how best to handle it? Whether SHA-256 was the best solution?
I don’t see anywhere that it talks about alternate proposals or why they might have been better. Why the particular suggestions here were rejected.
This seems like a bunch of Monday morning quarterbacking.
Edit: I'm not suggesting you should have written it any differently. I think you made the right choice to be a bit provocative because it grabs the attention that's needed.
(No clue if it's applicable here, I'm not aware of this case, but I believe that's the reference if it helps :)
Edit : exact quote, as Arthur's house is about to be demolished for a highway bypass:
"But the plans were on display…”
“On display? I eventually had to go down to the cellar to find them.”
“That’s the display department.”
“With a flashlight.”
“Ah, well, the lights had probably gone.”
“So had the stairs.”
“But look, you found the notice, didn’t you?”
“Yes,” said Arthur, “yes I did. It was on display in the bottom of a locked filing cabinet stuck in a disused lavatory with a sign on the door saying ‘Beware of the Leopard.
In what way has git's discussion of their move been hidden away in a metaphorical basement?
Additionally, as someone who read Hitchiker's guide to the Galaxy ~3 decades ago, I managed to suspect it was a reference to it, even if the reference wasn't the most applicable
Mailing lists are basements in 2026
The stewards of Git are going to do whatever they want, and there is nothing you can do about it if you don't have the clout to create a fork that takes the lead.
No amount of discussion will do anything because they've already decided that their view of the situation is correct. Git hashes are not just content identification but a digital certificate mechanism, and their collision resistance is a grave issue that must be fixed, the end.
You will be browbeaten in any discussion; it's not worth the energy in a world replete with issues.
I came in expecting to disagree strongly with the article, but ended up agreeing more than I didn't (though I still don't 100% agree, as collisions are still an issue for mirrors). I find the concept of multiple hashes per commit quite interesting. It would allow mixing hashes in one repo, wouldn't break submodules, and tooling could be used to reject commits without any secure hashes for a gradual transition (like enforcing signed commits/tags).
Also GitHub should auto-detect the format on the very first “git push” and handle it then. If they actually require config to be set appropriately during the initial “new repository” dialog that’s just bad ux.
As for this being a looming disaster for industry… it’s inconvenient. And lots of things will break. But we will survive and come out the other side I’m certain. We converted from the Julian calendar to the Gregorian calendar 500 years ago. Surely we can handle this.
thought "costly" in the title and "incomprehensibly expensive" in the subheader meant this piece would discuss how much less performant sha-256 is on modern machines, but didn't see anything. isn't there hardware acceleration? how much worse is it?
I just sent a patch series to the list that enables sha1dc to be accelerated on modern CPU architectures to close to normal SHA1 speeds, but since it was ported from a Rust project by an agent, it will never be applied.
https://lore.kernel.org/git/20260929112544.86511-1-scott@git...
As migrations go, it's reading as simple to me. You'll just have to backpoint the commit signatures. I must assume there's a backwards compatible reference for them in git 3, right?
Or drop them and reference the old structure in a dire pinch.
Once you rehash the entire repo, every single one of those external references will be broken. Because no, there's no support for looking up old hash -> new hash or the reverse.
But I also don't think that switching to SHA-256 must be painful. A git2->git3 converted repo could just store all the past hashes, so existing links don't break.
This even would make the SHA-1 git objects collision resistant when stored on a trusted server, even with untrusted clients. (Exercise left to the reader.)
If you need to certify the authenticity of some code, and you've decided that a Git hash of any kind is going to be your certificate, you have a problem between keyboard and chair which is not fixable by stronger hashes in Git.
I don't want instability and churn in tooling.
1. Collisions aren't as bad as preimage attacks
2. Even if you made a file-with-malicious-hash, how would you get people to pull it?
3. Other attacks are a bigger problem (social engineering)
(2) is laughable in a world with github. It's common for unknown people to submit pull requests to code bases, and for those changes to be reviewed and merged. For example, as part of reviewing pull requests, I have `git fetch`'d proposed changes to my local machine to check behavior on some additional test cases. "If you fetch it you're fucked" is unacceptable as a security boundary.
(1) and (3) are just tu-quoque arguments about other attacks being worse. The relevant question isn't how bad other attacks are, it's how bad this attack is.
The fundamental problem with collisions is that software often assumes they can't happen (or is not tested against them). Thus collisions can trigger bugs, or otherwise cause surprising behavior. For example, webkit figured the colliding PDFs demonstrating a sha1 collision would be excellent for unit tests, so they merged the PDFs into their SVN repo... which completely fucked it [1]. I don't know the exact internals of git so I can't comment on how you would get surprising things to happen, but "oops the file you merged was different than the file you reviewed" and "oops the repository got corrupted" seem entirely plausible.
[1]: https://www.reddit.com/r/programming/comments/5vyhy2/webkit_...
(1/3 counter) is not what I argued. I argued from the worst-case position that collision and preimages were theoretically cheap and fast. Even in that case, I feel my arguments hold.
The main issue here is that you assume you can replace an existing object with a replaced one, which you cannot. Not only that, but in all known cases, the sha1dc variant of SHA1 that Git uses will even _tell_ you that someone tried to do this, which singles out the source quickly.
Best I can tell, all a forced collision would do is let someone who already has control of a repo modify the history in a far from plausibly deniable way. Which in practical terms, they already could do simply by replacing the whole thing, because who's out here using git hashes as a security tool? Every pinning I've ever seen has been to tags (which can be modified at will), or hashes of the actual payload (which doesn't need to be the same as what git uses).
Among others, dependency management in Rust [1] and Python [2] sometimes uses references that work similar to https://github.com/rust-lang/rust/commit/ec999ed [3] to suggest one particular version of the project, authored by the specified maintainer.
Unfortunately, it means neither, unless you pushed it. The hash points to whatever the first person uploading it to github submitted. And the author/org name in the URL is window dressing: all the objects go in one big bucket regardless of push permission to one particular fork (because why wouldn't they - today, collisions are believed to be recognizable because the cheapest way to craft them results in clear tells).
[1]: https://doc.rust-lang.org/cargo/reference/specifying-depende...
[2]: https://pip.pypa.io/en/stable/topics/vcs-support/#git
[3]: N.B. the "This commit does not belong to any branch on this repository, and may belong to a fork outside of the repository." warning Github has started to add to URLs like that.
SHA-256 is considered quantum safe by the NIST and is left out of PQC migration guidance entirely.
Things like this have a tendency to to be revised as time passes.
What I'm missing in the article is whether any Git server accepts replacing a SHA-1 identified object it already has. If it doesn't, then the distribution trust discussed holds, and keeping SHA-1 seems fine. Adding additional signatures seems fine for those who need transitive trust.
Plain git init could fail with a diagnostic: informing to use one of the two aliases or an option.
What people don't want is making git repos SHA-256 by accident and finding out later that they made repos not compatible with older git.
I would write this to the mailing list, but I thought a conversation that includes people outside that list is more interesting to me. Ultimately I'm not sure if I'm dumb about this or the whistle blower that's willing to actually say "maybe this isn't the right call"
a) it's relatively fast and impossible in a practical sense for two different files to accidentally hash to the same value.
That is silly. We are not worried about accidentally triggering. We are worried about intentional triggers.
I dont know why people always bring this up for hashing. In any other context it would be considered silly. If someone said, the chance of triggering a buffer overflow by accident is low, we would call that silly as we aren't worried about accidental triggers.
b) second pre-image vs collision. In a world of open source where we accept commits from randoms on the internet, i think collisions are just as relavent as second pre-image.
b) so, you agree with the blog? "So, any realistic interesting attack vector therefore relies on a collision attack,"
My reading of the blog is that they are dismissive of collision attacks. In context of git, i disagree. I think there are plausible attack scenarios involving collisions, or at least, just as plausible as second pre-image.
If you mean do i agree with the blog that impossible attacks aren't possible? well yes obviously, but i think that goes without saying.
There's lots of tools that refer to git commit IDs. Some of those tools may even hardcode a commit ID to be 40 hex digits long. The fact that these tools are external also means that "oh, just rewrite the commit messages or code to refer to the new IDs" isn't feasible. The only way to not break the world is to let people refer to existing commits with their SHA-1 hashes in perpetuity, and it doesn't sound like git is set up to allow this in any way, which means that existing repositories have to stay SHA-1 in perpetuity and that will cause fun down the line if you start having to make SHA-1 and SHA-256 repositories.
Changing from master to main is a one-off change. It might require changing your scripts once to refer to 'origin/main' instead of 'origin/master', but other than that, there is essentially nothing more that needs to be done, there is no risk to historical artifacts that needs to be mitigated.
like gitc0ffee?
- SLSA and Provenance or SBOM data in the supply chain security that uses commit hash. All the previous images are now pointing to a non-existing commit
- All the documentation and tooling as the article calls out
- All your traceability links from your project tool to your git repo, they will lose all the past data as it will be dead links
So I hope there IS NOT a migration path for in-place replacement!
So given that, I'm more interested in the arguments for why migrating to SHA-256 is problematic.
The biggest issue I see, after skimming over it, is the submodule breakage for new projects trying to link to old projects. This seems solvable frankly, but is the only serious issue I see. Everything else will be worked out as software is updated IMO.
The arguments make sense, but how ironclad are they? How confident are you that some clever black hat won’t figure out a way to take advantage of it?
This is one of the most widely used programs in the world. Let’s close the hole.
Submodules is a legitimate argument against this, though I don't know how widely this feature is actually used, and similar to the arguments in favor of switching the default branch from master to main, this is simply a setting which can be changed.
I do like the idea of commits having both hashes, and am surprised that idea has not been explored further.
Generally though, I think the author's strongest argument is simply that the change isn't strictly "needed", and all the other issues presented aren't the strongest arguments against change.
Ecosystems like Yocto are built around having meta layers as submodules. And, despite the usability flaws of submodules, it works really well.
I also use submodules to include dependencies into C++ projects a lot. It works fine.
How does it work with MRs, can I submit an MR which consists of changing the referenced SHA (and have it not show up as changes to every file in the referenced repo)?
The trade-offs are relatively obvious. It'd be a poor option for Yocto, but is a better option for most corporate repos.
I really don't get the hate. They're not hard to work with. Just a bit shitty UX but if you're using Git you're used to that already.
Someone started this FUD a long time ago and it has worked. Instead of using an elegant mechanism, project have built inelegant wrappers on top of git like go.mod which are actual mistakes.
I say this as a person who strongly dislikes many many aspects of Unix and Linux due to bad design and terrible UX. Git has a better design than any Unix program you get.
Submodules work okay. It is just Git LFS but for Git repos. Get over it.
Compare the UX of go mod with git submodules. One is easy and the other is about as fun as having teeth extracted.
git’s UX has never been its strong point. But submodules takes that pain to a whole new level.
Prior to SHA1 we had MD5, a decade earlier. MD5 collision attacks had already been widely documented and known. It was the most obvious thing on Earth that this would happen to SHA1 too. Apparently, Linus never realized there was a need for cryptographic security and that the hash was purely internal.
Here's what I honestly think was a factor. I think C programmers fell in love with the implementation that you could throw around a fixed hash record on the stack. It's incredibly efficient. But it's an efficiency that doesn't really matter because as soon as you read from or write to a disk or a network or even memory, any cost saving is completely gone.
More than a decade ago, some people wrote a Java implementation of git (jgit?) and despite all their optimizations, it was (IIRC) only half as fast as C git. It is of course because Java at the time had no concept of stack values for non-primitive types so couldn't compete. Personally, I was impressed: only half the speed? That's pretty good.
For something that's only 20 years old, the Git SHA1 assumption is some of the worst technical debt we have in the modern era.
Here's another thought: when people make a lot of these programs, they often make the mistake of not separating the program version and the network protocol (or just the external API). So you end up with brittle client-server implementations where you have to upgrade both the client and the server at the same time because they lack a network abstraction.
The other end of the spectrum is video streaming where you have codex, container formats, transport protocols and so on.
What a mess.
Hahahahaha, that's some level of delusion
The alternative to making sha256 the default is to leave sha1 the default. Nobody changes to sha256. sha1 is broken in 10 years. Suddenly everyone has to switch all at once on the same day because it is a critical security issue, but github never implemented sha256 because they didn't have to. This would be a major problem.
This is very very easy to fix if you run into it.
1. Adopt git 3.0 if you can with sha256.
2. If you can't use sha256, set the config to put things back to sha1. Wherever you need to do this you probably already set dozens of ENV vars or settings, just add a new one.
Or write a 15 page analysis about how the above is so hard people will probably just find it catastrophic to even think about.
If you read the OP article, the entire point he's making is that this would never happen, because a hash algorithm being "broken" doesn't matter in practice, because true supply chain security has nothing to do with file hashes.
That said, it’s still a good idea to migrate to a more robust hashing algorithm. Defense in depth, etc. Just because it’s a difficult migration doesn’t mean it shouldn’t be done.
Who is "you" in the context of a distributed version control system? I think this is not just the plural you, but the unbounded you -- it's all people who not just interact with your project now, but who you hope may interact with it in the future. The question is what the cost is of committing a near-infinite population to this migration, not the cost of doing a single `brew update` on your personal machine, no?
Because a lot of work was done to prepare and fix potential issues.
I thought the master to main thing was bad enough but this is going to suck. And just like the master rename it achieves basically nothing.
What is it about these projects that attracts people who just want to change things for the sake of it? Real engineering means coming up with a solution for backwards compatibility. This is just irresponsible and, frankly, a fuck you to everyone who will be affected by this.
« What is it about these projects » — Maybe that they're “at the forefront.”
> We could be using MD5 and it would honestly probably be just fine.
https://lore.kernel.org/git/Pine.LNX.4.58.0504291221250.1890...
As linked by another commenter in this thread, Linus worked out years ago that even if someone inserted a malicious object into the kernel repo, it would at best be a nuisance and not a major concern.
Anyway, I think it'll be the same. Tools will support SHA-256 quickly and we might have a migration program that converts SHA-1 repo to SHA-256 repo.
The only problem is that git (and related tools) will get twice as big...
Besides the possible implementation/deployment issues they will or will not face with this update, I can empathize with the idea that of not wanting to have a possible vector of attack in your system. Particularly today with AI being able to find novel exploits, I could see a future where a vulnerable hashing system leads to a malicious injection attack.
The author argues that "If I wanted to get untrusted code into Android, it's so much simpler to bribe or convince the maintainer of a popular downstream project" which is a really a red herring in this matter since that is literally a completely different issue that obviously no software update could ever fix.
Nonetheless I do agree with him in regards of how complicated and messy this whole process will be. Crypto migrations have been historically difficult, expensive and overall ugly, but not impossible...
https://nvlpubs.nist.gov/nistpubs/gcr/2018/NIST.GCR.18-017.p... page 58 for instance.