Extracts textual content from images and scans using client-side Tesseract.js with per-word confidence scoring.