This article is full of mistakes and misleading claims:
1) It's claiming SHA1 insecurity is theoretical, while SHAttered from 2017 was specifically a pratical proof of concept. The only reason Git wasn't affected, is because they didn't bother bruteforcing a git-blob prefix.
2) It's claiming collision attacks don't matter, only second-preimage attacks do. This is incorrect, collision attacks are enough for code-smuggling problems, when two repositories are on the same git commit (verified by the full commit hash), yet contain different code in their git checkout.
3) The Linus quote "The real security is in distribution" is arguing that "git's content-addressed system should not be used to address content". It's arguing that, in case of curl|sh, you shouldn't use a sha256sum-gate to pin the content to something you've reviewed, you should instead ensure curl is fetching from an https server.
1) I link to the SHAttered paper, as well as Shambles. Git projects were not affected because it is an inefficient attack vector. I say it's impractical to exploit, which I think everyone agrees with.
2) I specifically argue that even if both attacks were practical and cheap, it's still not the problem we should be focusing on.
3) Have you read this email (that I linked to)? It is almost the same general message (20 years ago) that this blog post is. It literally goes though a theoretical object replacement attack and how dumb this scenario is and so SHA-1 is fine.
I don't understand why Git is not making the SHA-1 and SHA-256 modes far more compatible with each other.
SHA1-hashed objects should be able to refer to SHA-256-hashed objects, although this seems somewhat pointless.
But SHA-256-hashed objects should also be able to refer to SHA1-hashed objects, with a major caveat: if those objects themselves are part of a collision pair, then there is a genuine problem. But this is avoidable! Suppose that Linux decided to migrate to SHA-256. The upstream project could choose a pair of dates, say January 1 2027 and March 1 2027. Up to the first date, maintainers would be welcome to submit hashes of objects that are not yet in the repo but that they think they might submit later on, and, on that date, the upstream tree would finalize the list of these objects and reference it in the repo (with a new mechanism for this purpose). Effective the second date, the repo would start publishing SHA-256 commits and would never again accept a SHA1-hashed object that was not in the repo at the cutoff date or referenced as part of the Jan 1 block.
And now it would be impossible to get a new SHA1 collision in to the repo.
The only new git features needed would be:
a) actual compatibility so that a SHA-256-hashed object could reference a SHA1-hashed object
b) a new object type that's a list of allowed SHA1 hashes (or probably a tree of them) that is itself hashed with SHA-256 and a mechanism to link to one of these from a commit
c) a policy mechanism to set a repo to only allow SHA1-hashed-objects that a reachable from a preconfigured SHA-256-hashed commit
I skimmed the video, and I didn't quite catch that. Near the end of the video, however, she did mention that interop is in the works[0].
In any case, even if Git 3.0 were completely incompatible, it would suck, but it's not the end of the world. You just treat it as if you were migrating from one SCM system to another. CVS -> SVN -> Perforce -> Git -> Git 3.0 -> [...] been-there-done-that. This is something that both open-source and commercial projects have had to deal with over the years.
Or maybe it would be a repeat of Python 2.x -> 3.x. ¯\_(ツ)_/¯ With AI assistance, hopefully porting the tooling over may go a lot quicker and smoother.
One of my favorite fun facts about Fossil SCM (another source control by the devs of sqlite) is that they patched their use of SHA1 6 days after the shattered attack was published:
"Both Fossil and Git started out using only SHA1 hashes. But when the SHAttered attack against SHA1 was published on 2017-02-23, the need to migrate to a stronger hash algorithm was recognized. Fossil added the ability to use SHA3-256 as an alternative on 2017-03-01 (six days after the SHAttered attack was first published). SHA3-256 is now the default for all new repositories and check-ins in Fossil, though older check-ins that occurred prior to SHAttered can still use their original SHA1 hash. Hence, no repositories had to be rebuilt and no hyperlinks were broken."
To me it's so interesting watching in realtime Git is still battling with this decision and for Fossil it was just another week of development.
That whole page is fun to read. Another fun fact somewhere else in the docs is that Fossil uses a grow-only set to store commits. They came up with this scheme some years before it was formalized by CRDTs!
I mean, there are two things here. One is how difficult it is to have a different hashing mechanism. Brian and other heroes in the Git core group have done amazing work to make this _technically_ possible on a repo level. To test some of my theories, I trivially implemented MD5 and an insanely dumb and easily breakable hash backend. It's not _hard_ to change the mechanism now. It's about the community.
Fossil isn't difficult to change not because it's technically harder for Git but because Git has a community and ecosystem that Fossil does not. The cost is not in the individual project for Git, the cost is because there is _so much_ in Git and this bifurcates everything.
So they’ve been talking about this for many years, planning, and finally announce when they’re going to switch the default.
So this is the right time to post that everything they’re doing is wrong? Did you engage in all the discussions about it and how best to handle it? Whether SHA-256 was the best solution?
I don’t see anywhere that it talks about alternate proposals or why they might have been better. Why the particular suggestions here were rejected.
This seems like a bunch of Monday morning quarterbacking.
I do mention this in like the first paragraph. I don't feel great about it, but I've listened to these issues for years now during contributor summits and Git Merge talks and while it's always seemed problematic, I thought they would come up with a good solution. This last Git Merge confirmed that it's close to the switch and not in any way solved or improved. I don't want to just go with it for groupthink reasons. I never thought it was a good idea and I have said that, but we have a last chance to rethink this, so I'm curious if I'm alone or in the silent majority.
Your argument is persuasive and well illustrated. I think the problem is the intro paragraphs come off as too certain of catastrophe which, when juxtaposed with your claim that "smarter people than me have been working on this", makes it sound like you don't actually believe they're smarter than you. The rest of your essay feels fair and not judgmental.
It's a hitchhiker guide to the galaxy reference, where sure, something is technically available but not clearly published and there are hoops even for those who know what they're looking for.
No clue if it's applicable here but that's the reference :)
Edit : exact quote, as Arthur's house is about to be demolished for a highway bypass:
"But the plans were on display…”
“On display? I eventually had to go down to the cellar to find them.”
“That’s the display department.”
“With a flashlight.”
“Ah, well, the lights had probably gone.”
“So had the stairs.”
“But look, you found the notice, didn’t you?”
“Yes,” said Arthur, “yes I did. It was on display in the bottom of a locked filing cabinet stuck in a disused lavatory with a sign on the door saying ‘Beware of the Leopard.
nofunsir is a stoichastic parrot, matching to a bit in Hitchhikers Guide to the Galaxy, wherin the protagonist should have known to protest a plan to demolish his home where plans where clearly documented in a hard to find place that they could not have known about. It is not a good pattern match, because git has been discussing this in public on documented mailing lists for years.
Once this starts being actual pain, we will each vibe the replacement index creator (git already supports replacement objects), for back-forth conversion, populated on pack and object indexing.
For massive perf and mem use damage. But oh well. And then we will wait for official version
Tools like git-filter-repo[1] support rewriting commit hashes in commit messages. git-filter-repo actually does it by default; see `--preserve-commit-hashes` in the manual[2].
>it will be an incomprehensibly expensive and ultimately valueless and avoidable global nightmare.
thought "costly" in the title and "incomprehensibly expensive" in the subheader meant this piece would discuss how much less performant sha-256 is on modern machines, but didn't see anything. isn't there hardware acceleration? how much worse is it?
Is it? There are many benefits in principle to IPv6, but if my ISP continues to assign me a single dynamic IP, those benefits are entirely moot for me.
Right, but just because you don't happen to have IPv6 right now, how does that remove the benefits for others to have IPv6? That's like saying having a faster CPU wouldn't mean faster performance, because I don't have that CPU yet.
The difference with IPv6 adoption is that the internet relies heavily on network effects: so long as some hosts only have an IPv4 address, you need an IPv4 address for full connectivity, but then if everyone has an IPv4 address anyway, there is no immediate need to migrate to IPv6.
(Yes us Hacker News users have plenty of use cases for IPv6, like self-hosting and peer-to-peer networking and so on; we are not the average user.)
This effect doesn't exist for the Git migration. Each repo can be updated independently; it doesn't affect users of other repositories, and most likely, the majority of devs will work on some SHA-1 repos and some SHA-256 repos with no issue.
If anything, I would compare it with the Python 2 to Python 3 migration, which was also painful, but succeeded eventually (despite being much less necessary in the first place).
There's a massive push right now from top down to have secure software supply chains. Google SBOM and SigStore. It's not an organic need but if you have government customers you don't have many options.
I thought this would be a snark but it's an extremely well put together argument against the "Hashmageddon".
If you're replacing the weakness of SHA-1 just by going to another algorithm, you better be prepared to go to the next one when sha256 collisions happen, and it doesn't sound like git's design would be easy to modify for this type of crypto agility.
I do like their proposal for using signatures to establish trust and allow swapping sha256 for whatever comes next.
Yes every repo is either one or the other but you fix that by rehashing the entire repo. Everyone can do this independently. It's entirely possible to maintain to identical repos in SHA1 and SHA256 mode but for the most part I suspect once updated people will simply pull down the new repo and use git 3.0 as a required version.
As migrations go, it's reading as simple to me. You'll just have to backpoint the commit signatures. I must assume there's a backwards compatible reference for them in git 3, right?
Or drop them and reference the old structure in a dire pinch.
Do you have any external references to any commits that matter, for example in your communication platforms (emails, Slack) or your bug tracker? Or, worse yet, in places where they aren't just text format references, but used for things like CI/CD caching decisions or security scans?
Once you rehash the entire repo, every single one of those external references will be broken. Because no, there's no support for looking up old hash -> new hash or the reverse.
Re-hash the entire repo as in rewriting all history? Hooo boy will that be a mess, I deal with things which reverence commits by hash in repos all the damn time. There are thousands of them in every Yocto project!
"The migration to the new format is simple; Just re-write everything in the new format, but also keep the old format around forever too since data is lost in the new format!"
Can someone more cyber-pilled than me explain what the actual risk with Git hashes being susceptible to collision attacks is? Obviously accidental collisions are problematic, but to my understanding the probability of that is still approximately zero.
Best I can tell, all a forced collision would do is let someone who already has control of a repo modify the history in a far from plausibly deniable way. Which in practical terms, they already could do simply by replacing the whole thing, because who's out here using git hashes as a security tool? Every pinning I've ever seen has been to tags (which can be modified at will), or hashes of the actual payload (which doesn't need to be the same as what git uses).
The post's argument that hash collisions are irrelevant in practice is not convincing at all. Basically they amount to:
1. Collisions aren't as bad as preimage attacks
2. Even if you made a file-with-malicious-hash, how would you get people to pull it?
3. Other attacks are a bigger problem (social engineering)
(2) is laughable in a world with github. It's common for unknown people to submit pull requests to code bases, and for those changes to be reviewed and merged. For example, as part of reviewing pull requests, I have `git fetch`'d proposed changes to my local machine to check behavior on some additional test cases. "If you fetch it you're fucked" is unacceptable as a security boundary.
(1) and (3) are just tu-quoque arguments about other attacks being worse. The relevant question isn't how bad other attacks are, it's how bad this attack is.
The fundamental problem with collisions is that software often assumes they can't happen (or is not tested against them). Thus collisions can trigger bugs, or otherwise cause surprising behavior. For example, webkit figured the colliding PDFs demonstrating a sha1 collision would be excellent for unit tests, so they merged the PDFs into their SVN repo... which completely fucked it [1]. I don't know the exact internals of git so I can't comment on how you would get surprising things to happen, but "oops the file you merged was different than the file you reviewed" and "oops the repository got corrupted" seem entirely plausible.
schacon: Really like the "Independent Tree Hash Headers" idea.
How difficult would this be to get this functionality into git?
Would it cause any breaking changes with older versions?
Have you discussed this with any git devs to see if they are open to adding it?
Ugh, I didn't know that SHA-1 submodules wouldn't be supported in SHA-256 repos. That changes the transition from painless to a major dumpster fire. Having to maintain converted forks, and use different hashes from upstream is going to be a mess.
A lot of replies here seem to be asserting that this "isn't that hard" without addressing the thing that makes it most hard: submodule compatibility and the breadth of tooling
They're the best way we have to reference other repositories from one repository. All other solutions don't have the benefit of being built in to git and having support built in to all git forges.
Ecosystems like Yocto are built around having meta layers as submodules. And, despite the usability flaws of submodules, it works really well.
I also use submodules to include dependencies into C++ projects a lot. It works fine.
git subtree and git subrepo are compatible with all git forges and don't require normal developers to install the extensions. Only the person/bot doing the occasional sync to the external repo has to install the extension. I prefer git subrepo for most (but not all) use cases.
What's the advantage to using git subtree or git subrepo instead of git submodules? I've never heard of this, what's the difference between them? If it's an extension, how do people without the extensions end up downloading the code from the other repos?
How does it work with MRs, can I submit an MR which consists of changing the referenced SHA (and have it not show up as changes to every file in the referenced repo)?
They work by copying one repo inside another and providing tools to copy/sync it back out again. It's not a link, it's a copy. It's almost the same as copying the files into your repo and git add'ing them, but there are accounting and tools to pull changes from the subrepo back to the external repo.
The trade-offs are relatively obvious. It'd be a poor option for Yocto, but is a better option for most corporate repos.
Oh, I didn't want to vendor another repo into mine, I just want to store a reference to it. I'll keep using submodules then, as they're easier to work with than tools like gclient and repo.
I really don't get the hate. They're not hard to work with. Just a bit shitty UX but if you're using Git you're used to that already.
Note that subtree and subrepo have the same SHA-1/SHA-256 incompatibility issue that submodules do, so this will be just as much of a trainwreck for them as well.
Someone started this FUD a long time ago and it has worked. Instead of using an elegant mechanism, project have built inelegant wrappers on top of git like go.mod which are actual mistakes.
The first reason the author lists for why this will be bad is only an "issue" on Git hosts that don't allow repo creation on push (which is brain dead of GitHub). Any other host, you push your new repo, and it will see the hashing algorithm, and receive the contents accordingly.
Submodules is a legitimate argument against this, though I don't know how widely this feature is actually used, and similar to the arguments in favor of switching the default branch from master to main, this is simply a setting which can be changed.
I do like the idea of commits having both hashes, and am surprised that idea has not been explored further.
Generally though, I think the author's strongest argument is simply that the change isn't strictly "needed", and all the other issues presented aren't the strongest arguments against change.
The alternative to making sha256 the default is to leave sha1 the default. Nobody changes to sha256. sha1 is broken in 10 years. Suddenly everyone has to switch all at once on the same day because it is a critical security issue, but github never implemented sha256 because they didn't have to. This would be a major problem.
This is very very easy to fix if you run into it.
1. Adopt git 3.0 if you can with sha256.
2. If you can't use sha256, set the config to put things back to sha1. Wherever you need to do this you probably already set dozens of ENV vars or settings, just add a new one.
Or write a 15 page analysis about how the above is so hard people will probably just find it catastrophic to even think about.
It's easy for _one person_ to fix. It's not easy for the entire git ecosystem as a whole. GitHub, large internal corporate git repos, CI/CD systems, projects with submodules, etc. The second half the article explains all of this.
It’s already broken, but even though it’s broken it’s hard to generate git collisions because of the repo metadata. It’s easy to generate (for instance) standalone PDFs with identical hashes, but doing this with git in a useful way is much harder.
That said, it’s still a good idea to migrate to a more robust hashing algorithm. Defense in depth, etc. Just because it’s a difficult migration doesn’t mean it shouldn’t be done.
> sha1 is broken in 10 years. Suddenly everyone has to switch all at once on the same day because it is a critical security issue
If you read the OP article, the entire point he's making is that this would never happen, because a hash algorithm being "broken" doesn't matter in practice, because true supply chain security has nothing to do with file hashes.
Who is "you" in the context of a distributed version control system? I think this is not just the plural you, but the unbounded you -- it's all people who not just interact with your project now, but who you hope may interact with it in the future. The question is what the cost is of committing a near-infinite population to this migration, not the cost of doing a single `brew update` on your personal machine, no?
You can certainly do this, as I said, this is Google's backup plan. But defaults matter. People will start running this and getting repos that are uselessly incompatible with other repos, tools, libraries and server instances. Having it as an option is one thing. Making it a default will cause a lot of pain for people who don't want to care about this.
No, my argument is that the change should not happen at all and nobody wants it and it gains the community very, very little but the default change is forcing it on everyone and most will be _entirely_ unaware - now having to solve problems that are difficult to understand. Defaults also matter when they are the wrong defaults.
For the record, Y2K was not fud. It was very real, in a long list of datetime problems that are to come. Further datetime problems are coming at scheduled dates.
It was definitely FUD. There was a real problem (date counters would roll over), but the impacts of it were so ridiculously overstated that it eclipsed any sane discussion of the issue. We had people at the time predicting that planes would literally fall out of the sky when Y2k hit, which was never a realistic possibility.
They address this very theoretical. In short: Yes. Which makes sense if you don't treat the hash as a form of security against malice, especially in the case of attacks that are already impractical, which is the entire thrust of the article.
[delayed]
This article is full of mistakes and misleading claims:
1) It's claiming SHA1 insecurity is theoretical, while SHAttered from 2017 was specifically a pratical proof of concept. The only reason Git wasn't affected, is because they didn't bother bruteforcing a git-blob prefix.
2) It's claiming collision attacks don't matter, only second-preimage attacks do. This is incorrect, collision attacks are enough for code-smuggling problems, when two repositories are on the same git commit (verified by the full commit hash), yet contain different code in their git checkout.
3) The Linus quote "The real security is in distribution" is arguing that "git's content-addressed system should not be used to address content". It's arguing that, in case of curl|sh, you shouldn't use a sha256sum-gate to pin the content to something you've reviewed, you should instead ensure curl is fetching from an https server.
1) I link to the SHAttered paper, as well as Shambles. Git projects were not affected because it is an inefficient attack vector. I say it's impractical to exploit, which I think everyone agrees with.
2) I specifically argue that even if both attacks were practical and cheap, it's still not the problem we should be focusing on.
3) Have you read this email (that I linked to)? It is almost the same general message (20 years ago) that this blog post is. It literally goes though a theoretical object replacement attack and how dumb this scenario is and so SHA-1 is fine.
https://lore.kernel.org/git/Pine.LNX.4.58.0504291221250.1890...
I don't understand why Git is not making the SHA-1 and SHA-256 modes far more compatible with each other.
SHA1-hashed objects should be able to refer to SHA-256-hashed objects, although this seems somewhat pointless.
But SHA-256-hashed objects should also be able to refer to SHA1-hashed objects, with a major caveat: if those objects themselves are part of a collision pair, then there is a genuine problem. But this is avoidable! Suppose that Linux decided to migrate to SHA-256. The upstream project could choose a pair of dates, say January 1 2027 and March 1 2027. Up to the first date, maintainers would be welcome to submit hashes of objects that are not yet in the repo but that they think they might submit later on, and, on that date, the upstream tree would finalize the list of these objects and reference it in the repo (with a new mechanism for this purpose). Effective the second date, the repo would start publishing SHA-256 commits and would never again accept a SHA1-hashed object that was not in the repo at the cutoff date or referenced as part of the Jan 1 block.
And now it would be impossible to get a new SHA1 collision in to the repo.
The only new git features needed would be:
a) actual compatibility so that a SHA-256-hashed object could reference a SHA1-hashed object
b) a new object type that's a list of allowed SHA1 hashes (or probably a tree of them) that is itself hashed with SHA-256 and a mechanism to link to one of these from a commit
c) a policy mechanism to set a repo to only allow SHA1-hashed-objects that a reachable from a preconfigured SHA-256-hashed commit
Emily's talk does a pretty good job of summarizing the issues with intermixing the hashes: https://youtu.be/eJJp0RE7cd4
I skimmed the video, and I didn't quite catch that. Near the end of the video, however, she did mention that interop is in the works[0].
In any case, even if Git 3.0 were completely incompatible, it would suck, but it's not the end of the world. You just treat it as if you were migrating from one SCM system to another. CVS -> SVN -> Perforce -> Git -> Git 3.0 -> [...] been-there-done-that. This is something that both open-source and commercial projects have had to deal with over the years.
Or maybe it would be a repeat of Python 2.x -> 3.x. ¯\_(ツ)_/¯ With AI assistance, hopefully porting the tooling over may go a lot quicker and smoother.
[0] https://www.youtube.com/watch?v=eJJp0RE7cd4&t=1134s
Her entire section 2 (starting at 6:12) is basically about why git does not and will never allow mixing of SHA1 and SHA256.
The interop discussed is using copybara as a copy tool to move data from SHA1 based repos to SHA256 based repos and vice versa.
One of my favorite fun facts about Fossil SCM (another source control by the devs of sqlite) is that they patched their use of SHA1 6 days after the shattered attack was published:
"Both Fossil and Git started out using only SHA1 hashes. But when the SHAttered attack against SHA1 was published on 2017-02-23, the need to migrate to a stronger hash algorithm was recognized. Fossil added the ability to use SHA3-256 as an alternative on 2017-03-01 (six days after the SHAttered attack was first published). SHA3-256 is now the default for all new repositories and check-ins in Fossil, though older check-ins that occurred prior to SHAttered can still use their original SHA1 hash. Hence, no repositories had to be rebuilt and no hyperlinks were broken."
https://fossil-scm.org/home/doc/trunk/www/hundredandone.md
To me it's so interesting watching in realtime Git is still battling with this decision and for Fossil it was just another week of development.
That whole page is fun to read. Another fun fact somewhere else in the docs is that Fossil uses a grow-only set to store commits. They came up with this scheme some years before it was formalized by CRDTs!
I mean, there are two things here. One is how difficult it is to have a different hashing mechanism. Brian and other heroes in the Git core group have done amazing work to make this _technically_ possible on a repo level. To test some of my theories, I trivially implemented MD5 and an insanely dumb and easily breakable hash backend. It's not _hard_ to change the mechanism now. It's about the community.
Fossil isn't difficult to change not because it's technically harder for Git but because Git has a community and ecosystem that Fossil does not. The cost is not in the individual project for Git, the cost is because there is _so much_ in Git and this bifurcates everything.
That's impressive. I suppose they had a more flexible architecture to make that change so fast.
Is there any writeup on why it was easy for them and not for git?
So they’ve been talking about this for many years, planning, and finally announce when they’re going to switch the default.
So this is the right time to post that everything they’re doing is wrong? Did you engage in all the discussions about it and how best to handle it? Whether SHA-256 was the best solution?
I don’t see anywhere that it talks about alternate proposals or why they might have been better. Why the particular suggestions here were rejected.
This seems like a bunch of Monday morning quarterbacking.
I do mention this in like the first paragraph. I don't feel great about it, but I've listened to these issues for years now during contributor summits and Git Merge talks and while it's always seemed problematic, I thought they would come up with a good solution. This last Git Merge confirmed that it's close to the switch and not in any way solved or improved. I don't want to just go with it for groupthink reasons. I never thought it was a good idea and I have said that, but we have a last chance to rethink this, so I'm curious if I'm alone or in the silent majority.
Your argument is persuasive and well illustrated. I think the problem is the intro paragraphs come off as too certain of catastrophe which, when juxtaposed with your claim that "smarter people than me have been working on this", makes it sound like you don't actually believe they're smarter than you. The rest of your essay feels fair and not judgmental.
The plans have been on display in a cellar. Beware of the leopard.
I don’t understand what this is supposed to mean.
It's a hitchhiker guide to the galaxy reference, where sure, something is technically available but not clearly published and there are hoops even for those who know what they're looking for.
No clue if it's applicable here but that's the reference :)
Edit : exact quote, as Arthur's house is about to be demolished for a highway bypass:
"But the plans were on display…”
“On display? I eventually had to go down to the cellar to find them.”
“That’s the display department.”
“With a flashlight.”
“Ah, well, the lights had probably gone.”
“So had the stairs.”
“But look, you found the notice, didn’t you?”
“Yes,” said Arthur, “yes I did. It was on display in the bottom of a locked filing cabinet stuck in a disused lavatory with a sign on the door saying ‘Beware of the Leopard.
It's not at all applicable.
nofunsir is a stoichastic parrot, matching to a bit in Hitchhikers Guide to the Galaxy, wherin the protagonist should have known to protest a plan to demolish his home where plans where clearly documented in a hard to find place that they could not have known about. It is not a good pattern match, because git has been discussing this in public on documented mailing lists for years.
This means, if you migrate your repo, every single commit message that contains text like: "please see commit <sha1>" will now be broken.
This will be a train wreck. I hope they don't release before adding compatibility modes to keep the existing sha1's around in the database.
Once this starts being actual pain, we will each vibe the replacement index creator (git already supports replacement objects), for back-forth conversion, populated on pack and object indexing.
For massive perf and mem use damage. But oh well. And then we will wait for official version
Tools like git-filter-repo[1] support rewriting commit hashes in commit messages. git-filter-repo actually does it by default; see `--preserve-commit-hashes` in the manual[2].
[1]: https://github.com/newren/git-filter-repo
[2]: https://htmlpreview.github.io/?https://github.com/newren/git...
Sure, but that's not going to rewrite Slack messages, emails, GitHub links, docs
Came to say something exactly like this... The old sha1 handle needs to be still available in the same way an HTTP 301 redirect would work.
Since SHA-1 is already broken (just expensive in terms of GPU-time), then the text "please see commit <sha1>" is also already broken.
You can't attack an existing normal commit.
But also collisions there aren't a big deal. People will cite short hashes when referring to things and that's not "broken".
>it will be an incomprehensibly expensive and ultimately valueless and avoidable global nightmare.
thought "costly" in the title and "incomprehensibly expensive" in the subheader meant this piece would discuss how much less performant sha-256 is on modern machines, but didn't see anything. isn't there hardware acceleration? how much worse is it?
He means costly in terms of human effort and wasted time.
From what I'd read, SHA256 in git is showing every sign of being another IPv6. In particular:
- It's implemented in a non-backwards-compatible way
- The benefits over the older model are a bit nebulous
- There's a large amount of tooling that needs to catch up, and little sign that there is movement there
> - The benefits over the older model are a bit nebulous
This is far from the case with IPv6!
Is it? There are many benefits in principle to IPv6, but if my ISP continues to assign me a single dynamic IP, those benefits are entirely moot for me.
If your ISP is assigning you a single IPv6 address, they are doing IPv6 wrong. You should be getting your own /56.
Even a /64, while wrong, would be a significant improvement over IPv4.
Your ISP assigns you a single dynamic /128? Which ISP is that?
Is the argument that there are no benefits to IPv6 for anyone/society because your specific ISP messes it up?
Right, but just because you don't happen to have IPv6 right now, how does that remove the benefits for others to have IPv6? That's like saying having a faster CPU wouldn't mean faster performance, because I don't have that CPU yet.
The difference with IPv6 adoption is that the internet relies heavily on network effects: so long as some hosts only have an IPv4 address, you need an IPv4 address for full connectivity, but then if everyone has an IPv4 address anyway, there is no immediate need to migrate to IPv6.
(Yes us Hacker News users have plenty of use cases for IPv6, like self-hosting and peer-to-peer networking and so on; we are not the average user.)
This effect doesn't exist for the Git migration. Each repo can be updated independently; it doesn't affect users of other repositories, and most likely, the majority of devs will work on some SHA-1 repos and some SHA-256 repos with no issue.
If anything, I would compare it with the Python 2 to Python 3 migration, which was also painful, but succeeded eventually (despite being much less necessary in the first place).
There's a massive push right now from top down to have secure software supply chains. Google SBOM and SigStore. It's not an organic need but if you have government customers you don't have many options.
Ironically rewriting git history is a perfect opportunity for a supply chain attack.
I thought this would be a snark but it's an extremely well put together argument against the "Hashmageddon".
If you're replacing the weakness of SHA-1 just by going to another algorithm, you better be prepared to go to the next one when sha256 collisions happen, and it doesn't sound like git's design would be easy to modify for this type of crypto agility.
I do like their proposal for using signatures to establish trust and allow swapping sha256 for whatever comes next.
Perhaps not clear enough, but Scott was a cofounder of GitHub, so he knows a thing or two about git in the real world =)
Yes every repo is either one or the other but you fix that by rehashing the entire repo. Everyone can do this independently. It's entirely possible to maintain to identical repos in SHA1 and SHA256 mode but for the most part I suspect once updated people will simply pull down the new repo and use git 3.0 as a required version.
As migrations go, it's reading as simple to me. You'll just have to backpoint the commit signatures. I must assume there's a backwards compatible reference for them in git 3, right?
Or drop them and reference the old structure in a dire pinch.
Do you have any external references to any commits that matter, for example in your communication platforms (emails, Slack) or your bug tracker? Or, worse yet, in places where they aren't just text format references, but used for things like CI/CD caching decisions or security scans?
Once you rehash the entire repo, every single one of those external references will be broken. Because no, there's no support for looking up old hash -> new hash or the reverse.
Re-hash the entire repo as in rewriting all history? Hooo boy will that be a mess, I deal with things which reverence commits by hash in repos all the damn time. There are thousands of them in every Yocto project!
Yes, this would cause big issues for Nix based build systems or any others that reference commits by hash.
"The migration to the new format is simple; Just re-write everything in the new format, but also keep the old format around forever too since data is lost in the new format!"
Can someone more cyber-pilled than me explain what the actual risk with Git hashes being susceptible to collision attacks is? Obviously accidental collisions are problematic, but to my understanding the probability of that is still approximately zero.
Best I can tell, all a forced collision would do is let someone who already has control of a repo modify the history in a far from plausibly deniable way. Which in practical terms, they already could do simply by replacing the whole thing, because who's out here using git hashes as a security tool? Every pinning I've ever seen has been to tags (which can be modified at will), or hashes of the actual payload (which doesn't need to be the same as what git uses).
I see commit hashes used all the time, like in yocto recipes for example
The post's argument that hash collisions are irrelevant in practice is not convincing at all. Basically they amount to:
1. Collisions aren't as bad as preimage attacks
2. Even if you made a file-with-malicious-hash, how would you get people to pull it?
3. Other attacks are a bigger problem (social engineering)
(2) is laughable in a world with github. It's common for unknown people to submit pull requests to code bases, and for those changes to be reviewed and merged. For example, as part of reviewing pull requests, I have `git fetch`'d proposed changes to my local machine to check behavior on some additional test cases. "If you fetch it you're fucked" is unacceptable as a security boundary.
(1) and (3) are just tu-quoque arguments about other attacks being worse. The relevant question isn't how bad other attacks are, it's how bad this attack is.
The fundamental problem with collisions is that software often assumes they can't happen (or is not tested against them). Thus collisions can trigger bugs, or otherwise cause surprising behavior. For example, webkit figured the colliding PDFs demonstrating a sha1 collision would be excellent for unit tests, so they merged the PDFs into their SVN repo... which completely fucked it [1]. I don't know the exact internals of git so I can't comment on how you would get surprising things to happen, but "oops the file you merged was different than the file you reviewed" and "oops the repository got corrupted" seem entirely plausible.
[1]: https://www.reddit.com/r/programming/comments/5vyhy2/webkit_...
couldn't github reject a push that contains an existing hash in the repo?
schacon: Really like the "Independent Tree Hash Headers" idea. How difficult would this be to get this functionality into git? Would it cause any breaking changes with older versions? Have you discussed this with any git devs to see if they are open to adding it?
> We can go through years of this SHA-1 to SHA-256 migration and then quantum computers break 256 and we're back in the same stupid boat again.
SHA-256 is considered quantum safe by the NIST and is left out of PQC migration guidance entirely.
For now. Turns out the pigeon hole principle still holds.
Ugh, I didn't know that SHA-1 submodules wouldn't be supported in SHA-256 repos. That changes the transition from painless to a major dumpster fire. Having to maintain converted forks, and use different hashes from upstream is going to be a mess.
git submodules has always been a major dumpster fire. You’re honestly better avoiding regardless of SHA-256 incompatibilities
The xz backdoor is basically his point in practice — that was a maintainer-trust compromise, not a hash collision.
A lot of replies here seem to be asserting that this "isn't that hard" without addressing the thing that makes it most hard: submodule compatibility and the breadth of tooling
Submodules are a mistake.
They're the best way we have to reference other repositories from one repository. All other solutions don't have the benefit of being built in to git and having support built in to all git forges.
Ecosystems like Yocto are built around having meta layers as submodules. And, despite the usability flaws of submodules, it works really well.
I also use submodules to include dependencies into C++ projects a lot. It works fine.
git subtree and git subrepo are compatible with all git forges and don't require normal developers to install the extensions. Only the person/bot doing the occasional sync to the external repo has to install the extension. I prefer git subrepo for most (but not all) use cases.
What's the advantage to using git subtree or git subrepo instead of git submodules? I've never heard of this, what's the difference between them? If it's an extension, how do people without the extensions end up downloading the code from the other repos?
How does it work with MRs, can I submit an MR which consists of changing the referenced SHA (and have it not show up as changes to every file in the referenced repo)?
They work by copying one repo inside another and providing tools to copy/sync it back out again. It's not a link, it's a copy. It's almost the same as copying the files into your repo and git add'ing them, but there are accounting and tools to pull changes from the subrepo back to the external repo.
The trade-offs are relatively obvious. It'd be a poor option for Yocto, but is a better option for most corporate repos.
Oh, I didn't want to vendor another repo into mine, I just want to store a reference to it. I'll keep using submodules then, as they're easier to work with than tools like gclient and repo.
I really don't get the hate. They're not hard to work with. Just a bit shitty UX but if you're using Git you're used to that already.
Note that subtree and subrepo have the same SHA-1/SHA-256 incompatibility issue that submodules do, so this will be just as much of a trainwreck for them as well.
> Submodules are a mistake.
Someone started this FUD a long time ago and it has worked. Instead of using an elegant mechanism, project have built inelegant wrappers on top of git like go.mod which are actual mistakes.
I’m not the author but I agree with their opinion.
Compare the UX of go mod with git submodules. One is easy and the other is about as fun as having teeth extracted.
git’s UX has never been its strong point. But submodules takes that pain to a whole new level.
I find them extremely useful. It gives me a monorepo experience in repos that otherwise aren't/can't exist as one monorepo for various reasons.
I run a small git/jj forge and for us it's already painful dealing with this. Can't imagine how GitHub is going to handle it.
The first reason the author lists for why this will be bad is only an "issue" on Git hosts that don't allow repo creation on push (which is brain dead of GitHub). Any other host, you push your new repo, and it will see the hashing algorithm, and receive the contents accordingly.
Submodules is a legitimate argument against this, though I don't know how widely this feature is actually used, and similar to the arguments in favor of switching the default branch from master to main, this is simply a setting which can be changed.
I do like the idea of commits having both hashes, and am surprised that idea has not been explored further.
Generally though, I think the author's strongest argument is simply that the change isn't strictly "needed", and all the other issues presented aren't the strongest arguments against change.
This seems like Y2K fud.
The alternative to making sha256 the default is to leave sha1 the default. Nobody changes to sha256. sha1 is broken in 10 years. Suddenly everyone has to switch all at once on the same day because it is a critical security issue, but github never implemented sha256 because they didn't have to. This would be a major problem.
This is very very easy to fix if you run into it.
1. Adopt git 3.0 if you can with sha256.
2. If you can't use sha256, set the config to put things back to sha1. Wherever you need to do this you probably already set dozens of ENV vars or settings, just add a new one.
Or write a 15 page analysis about how the above is so hard people will probably just find it catastrophic to even think about.
It's easy for _one person_ to fix. It's not easy for the entire git ecosystem as a whole. GitHub, large internal corporate git repos, CI/CD systems, projects with submodules, etc. The second half the article explains all of this.
It’s already broken, but even though it’s broken it’s hard to generate git collisions because of the repo metadata. It’s easy to generate (for instance) standalone PDFs with identical hashes, but doing this with git in a useful way is much harder.
That said, it’s still a good idea to migrate to a more robust hashing algorithm. Defense in depth, etc. Just because it’s a difficult migration doesn’t mean it shouldn’t be done.
> sha1 is broken in 10 years. Suddenly everyone has to switch all at once on the same day because it is a critical security issue
If you read the OP article, the entire point he's making is that this would never happen, because a hash algorithm being "broken" doesn't matter in practice, because true supply chain security has nothing to do with file hashes.
> Adopt git 3.0 if you can with sha256.
Who is "you" in the context of a distributed version control system? I think this is not just the plural you, but the unbounded you -- it's all people who not just interact with your project now, but who you hope may interact with it in the future. The question is what the cost is of committing a near-infinite population to this migration, not the cost of doing a single `brew update` on your personal machine, no?
Just clone the repo again. Jesus Christ, you act like the simplest thing in the world is some kind of insurmountable challenge.
You can certainly do this, as I said, this is Google's backup plan. But defaults matter. People will start running this and getting repos that are uselessly incompatible with other repos, tools, libraries and server instances. Having it as an option is one thing. Making it a default will cause a lot of pain for people who don't want to care about this.
"defaults matter" is an argument for this change, not against it.
No, my argument is that the change should not happen at all and nobody wants it and it gains the community very, very little but the default change is forcing it on everyone and most will be _entirely_ unaware - now having to solve problems that are difficult to understand. Defaults also matter when they are the wrong defaults.
This is not difficult to understand. It's very easy to understand.
For the record, Y2K was not fud. It was very real, in a long list of datetime problems that are to come. Further datetime problems are coming at scheduled dates.
It was definitely FUD. There was a real problem (date counters would roll over), but the impacts of it were so ridiculously overstated that it eclipsed any sane discussion of the issue. We had people at the time predicting that planes would literally fall out of the sky when Y2k hit, which was never a realistic possibility.
> which was never a realistic possibility.
Because a lot of work was done to prepare and fix potential issues.
…
What does OpenAI have anything to do with this?
Would the author feel the same if git had used MD5 instead of SHA-1?
I do actually literally write in this that if it was MD5 it also would not be a problem.
Fair cop.
They address this very theoretical. In short: Yes. Which makes sense if you don't treat the hash as a form of security against malice, especially in the case of attacks that are already impractical, which is the entire thrust of the article.