original author README: http://extractomatic.tomtaylor.co.uk/
Web service for HTML content extraction.