Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Yep! Glad you’re enjoying it. Here are my share of the SE productions: https://www.robinwhittleton.com/books/

Biggest so far (and biggest on SE at 1.2m words) has been Samuel Pepys’ Diary: https://standardebooks.org/ebooks/samuel-pepys/the-diary



Whoa, SE looks incredible. Thanks for your work on that!

Are there any plans to support non-English ebooks as well?

Edit: Regarding non-English ebooks, I was thinking about books like "Der Vogelflug als Grundlage der Fliegekunst" (Birdflight As The Basis Of Aviation) by Otto Lilienthal (https://www.gutenberg.org/ebooks/54565). It's a fantastic book with nice hand-drawn illustrations, and it would deserve being presented/typeset in a beautiful way. Currently available ebooks are mostly bad.


Sorry, no. Like I said in another comment, a large part of what makes SE good is our manual of style, which is really just applicable to English works. If you want to use the toolchain to produce a novel in another language feel free (removing our logo and name, of course).


Thank you for your efforts!

That book is more than 2600 pages! How long did that take?


Nearly two years, but I took a bunch of time out to work on smaller pieces. The biggest problem was that Gutenberg had only transcribed about a third of the footnotes, so I spent a long time adding the remaining ones back in.


Great work! Is there any link between this and the Pepys Diary site at https://www.pepysdiary.com (also an awesome resource)?


I had a chat with Phil after I finished, and he sent over a list of transcription problems with Gutenberg that I’d missed which was very kind. But apart from that, no direct collaboration.


So you were transcribing a diary written in a shorthand? Or from a print version?


From scanned original sources. I just opened a random page at archive.org to give you an example:

https://archive.org/details/diaryofsamuelpep01pepy/page/228/...


Do you use OCR? Sorry for the million questions.


No problem! Yes, archive.org has OCRed all the text, but it needs a lot of finessing to bring up the quality to an acceptable standard.

Typically, Standard Ebooks will try to avoid OCR unless it’s something that’s missing from Gutenberg.


Thank you for your time!

(If you started some sort of AMA here I'm sure it would be successful.)




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: