First, where do I find the most popular one million 12-character passwords that use A-Z/a-z/0-9?
Second, the top 1M of 6/7/8 character passwords have a statistically higher probably of being the correct password than the top 1M of 62^12 because their distribution function as a percentage of possible instances is narrower. Put another way: the top 7 char password might be used 10 million times, and the 2nd most popular might be 9million, etc. With a 12 char password, the top 1 by nature appears much less frequency because the total possible space is larger and people have to think a little more. Entropy when that has nothing to do with inverse distributions.
Put another way, look at the frequency of the top-1 password as a sum of all instances of all 1M passwords. It is much larger % for smaller passwords than larger. It is much less likely that the top 1M 12-char passwords are as likely to be successful as the top 1M 6/7/8-char passwords.
I just did this for the top 1000 6/7/8 chhar passwords, the top 6 digit password "123456" represents 1.3% of all 1000 6-chars, where as "password" is 0.9% of all 1000 8-char.
> Second, the top 1M of 6/7/8 character passwords have a statistically higher probably of being the correct password than the top 1M of 62^12 because their distribution function as a percentage of possible instances is narrower.
Not how this works. This would only be true if you were picking passwords randomly out of all possible ones and the attacker knew how long the password. For example, "password" is a more common password than "cat" is (both are terrible of course). It doesn't matter that "password" is longer than "cat".
If you knew how many characters long the password was it might be a different story. But you don't have that information. 012345 being a larger percentage of 6 character passwords than "password" of 8 character passwords means nothing if you don't know what length the password is. After all, 100% of passwords "af64nh8" are "af64nh8", but that means nothing unless you know the password you are attacking falls into that class.
There's been tons of dumps posted. Normally I see the download or magnet links on Reddit or HN. Even with hashes people have broken 69% [0] to 85% [1] of the accounts. There's mailing list or forums for various related tools (cracking tools, gpu cracking, have I been pwned, etc.) that discuss these kinds of things, including where to get the passwords and/or hashes.
One of the summaries claimed that 3% of the cracked passwords were 12 characters long (some 1.9M). 48% of the cracked passwords were lower case and numbers.
Often a thread that mentions the compromise/leak also mentions the newest "batch" that includes the new entries. The last one I tracked was RockYou2021.txt, some 100GB, around 10GB compressed. Just use your favorite egrep/perl/python to filter for whatever you want. I checked the a random torrent and found 25 people on it.
I agree on your comments on the distribution, but the average user has shockingly little entropy, and human password entropy doesn't scale with password length. I expect a decent chunk of the passwords are going to be whatever the user's default password was and either repeating or extending it in obvious ways. Even the simplest approach of testing 2 words or 3 words for a total 12-15 characters with the simple 3/E i/one and 7/L replacements would likely. Granted I'm not expecting the same 8 character 69-85% with 12 characters, but I also don't expect much significantly less, and I believe 33M accounts were stolen.
rockyou2021.txt.7z has 8459060239 passwords
rockyou2021.txt.7z has 835365123 that are 12 characters long
pwned-passwords has 847223402 unique password hashes
pwned-passwords has 5579399834 non-unique passwords hashes
I'm amused that each password was reused on average 6.5 times.
I suspect large fraction of the pwned-passwords hashes could be turned into passwords with the rockyou2021 list.
> First, where do I find the most popular one million 12-character passwords that use A-Z/a-z/0-9?
I'm sure it may not be easy to find this database, but it is probably not that difficult if you hang out in the type of circles that the people who cracked that site hang out in.
> Math...
This strikes me as nitpicking when one can access such powerful computing devices as to make the difference practically irrelevant. So make it 10million and wait a few more minutes, who cares?
How is it nitpicking? The distribution tail effectively goes to a frequency of zero. Increasing to 10M doesn't solve the problem because the space is so vast, by the time you get to a fraction that has the similar area under the curve as 1M for 6/7/8 chars, you've lost any compute advantage. We're not talking an order of magnitude, we're talking multiple. As much as I know y'all want to hate on any password manager that isn't your favorite, math doesn't care about what strikes you as nitpicking.
> by the time you get to a fraction that has the similar area under the curve as 1M for 6/7/8 chars,
You are (incorrectly) conflating entropy of a password with its length.
> math doesn't care about what strikes you as nitpicking.
Math actually cares a lot about the nitpicky details. Both in the sense that small things can have big effects, but also in the sense that things which sound big can also be irrelavent.
I don't use a password manager and am not familiar with the math involved; I was going off your 10million number and the post above you's 300,000 iterations/sec.
I'm trying to understand how Top-N frequency distribution flattens and becomes less useful as search space increases. It is a both a statistical and a psychological issue. If you can't help don't comment.
I think you might find that your life goes a little more smoothly if you weren't so abrasive, confrontational and dismissive. As a side affect, people you talk to won't end up feeling like shit because you possibly will treat them less terribly.
Second, the top 1M of 6/7/8 character passwords have a statistically higher probably of being the correct password than the top 1M of 62^12 because their distribution function as a percentage of possible instances is narrower. Put another way: the top 7 char password might be used 10 million times, and the 2nd most popular might be 9million, etc. With a 12 char password, the top 1 by nature appears much less frequency because the total possible space is larger and people have to think a little more. Entropy when that has nothing to do with inverse distributions.
Put another way, look at the frequency of the top-1 password as a sum of all instances of all 1M passwords. It is much larger % for smaller passwords than larger. It is much less likely that the top 1M 12-char passwords are as likely to be successful as the top 1M 6/7/8-char passwords.
I just did this for the top 1000 6/7/8 chhar passwords, the top 6 digit password "123456" represents 1.3% of all 1000 6-chars, where as "password" is 0.9% of all 1000 8-char.