Text extraction from MS Office and PDF

Vitali · July 24, 2010, 1:45pm

Hello,

I’m looking for libraries to do text extraction from MS Office and PDF
file formats. Also looking for libraries to do HTML rendering of
documents in the same formats. I know of couple of commercial
libraries from Oracle and Autonomy, but they only have C and/or Java
APIs. I also found this project POI Ruby Bindings.
Is there other open source alternatives, and/or alternatives with Ruby
bindings?

Thanks,
Vitali

Vitali · July 28, 2010, 3:34am

I am using the standard ‘spreadsheet’ library to load from excel
2003 .xls files with ruby.
it’s not pretty but it works