> >I'd vote for a GOOD OCR copy of everything on a CD, 30-40mb for a mostly text file is scandalous - there should be a 'quality' way to get an ocr scan from old hardcopy
>
> OCR reading documents are a labourous process. I once started to OCR
> my TNE rulebook. But after spending several hours to correct miss
> spelling like "lDiot" to read 1D10 I gave up. And I used at the time
> the best OCR software available and back then a quite powerful
> computer. The print and the typeface in the TNE rulebook isn't very
> clean and the rasterized sidebar boxes are hard to OCR. It takes less
> time typing them in by hand.
yep, i've tried "the Joy of OCR Correction" also... :|
and redoing the spelling was the least of it - you have to wonder where they get some of the gobbledygook characters from - and how they can mangle the paragraph order of a 'simple' one column document - and the tables... < "Oh, the Pain, the Pain..." >
if i could type worth a damn, it just might have been faster
anyone know how Project Gutenberg gets its OCR done??
do they have better software, or do they use multiple volunteers to scan small groups of pages - and then correct them separately?