Split Words in Pasted PDFs: Rejoining Hyphenated Lines
You copy a page from a justified PDF and the text is littered with fragments like split-word pairs: words the typesetter divided with hyphens to keep the right margin neat. Pasted into your editor, the hyphen and the line break survive, so search cannot find the whole word, spellcheck flags every fragment, and quoting the passage means manually stitching dozens of words back together. A single academic page can hold hundreds of these splits.
The PDF Hyphen Fixer rejoins those split words automatically. It only removes hyphens sitting at the very end of a line followed by a lowercase letter, which is the classic signature of PDF line-break hyphenation. Genuine hyphens inside words, like well-known or state-of-the-art, are never touched, and hyphens before capitalized words are preserved by default to protect proper names. Once words are whole again, feed the output into the line break fixer to repair the staircase of line breaks.
How does the tool know which hyphens to remove?
It looks for a hyphen at the end of a line followed by a lowercase letter starting the next line, which is the signature of PDF line-break hyphenation. Hyphens in the middle of a line, like in well-known, are never touched. Hyphens before capitalized words are left alone by default, since a capital signals a proper name or a compound.
This conservative design is deliberate. A tool that aggressively stripped every hyphen would corrupt real compounds and proper names, creating errors worse than the fragments it fixed. By acting only on the line-end-plus-lowercase pattern, the fixer stays on the safe side: when in doubt, it keeps the hyphen. The result note reports exactly how many words were rejoined, so you can compare input and output and verify each fix.
If your document legitimately splits many words before capitalized words, tick the also-rejoin-before-capitals option to catch those too, then review the output for proper names. For everything else, the default behavior is the right balance of thorough and safe.
Why do PDFs split words with hyphens in the first place?
PDFs are designed for fixed print layouts, so long words at the right margin get hyphenated to keep edges neat and justified. Copying preserves both the hyphen and the line break, leaving word fragments stranded on the next line. Justified text and narrow columns produce the most hyphenation, which is why academic papers and legal documents are the worst offenders.
Typesetting hyphenation follows dictionary rules, splitting words between syllables. When you read the printed page, your eye reassembles the word effortlessly. But copied text has no such reader; it is just characters, and a word split across a line break is a broken token to every tool that processes it. Search indexes the fragments separately, spellcheckers flag them, and text-to-speech reads the hyphen as a pause. None of these tools can guess the original word, which is why rejoining matters before any other processing.
The fixer reverses this typesetting artifact, restoring the original unbroken words for editing and searching. It is the essential first step in any PDF cleanup pipeline, because every later step, from line-break repair to spellchecking, depends on words being whole. Pair it with the invisible character detector when the PDF also carries soft hyphens, which are invisible Unicode characters rather than visible line-end hyphens.
What about soft hyphens and capitalized splits?
Soft hyphens are invisible Unicode characters, a different problem from line-end hyphens, so use the invisible character detector to find and remove them first. For splits where the next line starts with a capital letter, the tool keeps the hyphen by default to protect proper names. Enable the also-rejoin-before-capitals option only when your document splits many words before capitals.
The capital-letter rule deserves attention because it guards against a subtle error. Consider a line ending with a hyphenated prefix and the next line starting with a capitalized title like President. Joining them would be wrong; the hyphen is almost certainly a genuine compound. The default keeps it, and you can review these cases manually. Only when you know the document hyphenates ordinary words before capitals, common in some legal and technical texts, should you enable the aggressive option.
After rejoining, collapse any double spaces the repair left behind with the extra spaces remover, then continue to the line break fixer. The full chain, soft hyphens out, visible hyphens rejoined, lines repaired, spacing tidied, turns even the most damaged PDF paste into clean editable text.
How to use the PDF Hyphen Fixer in 4 steps
- Paste the raw PDF text. Use the original copy-paste with line breaks intact, since the tool reads line endings to find split words.
- Choose your options. The defaults protect real hyphens and capitalized words. Enable also-rejoin-before-capitals only if your document splits ordinary words before capitals.
- Run it and check the count. The result note tells you exactly how many words were rejoined. Compare with the input to verify each fix.
- Continue the cleanup chain. Feed the output into the line break fixer to repair the remaining line breaks, then tidy spacing.
5 practical tips for fixing PDF hyphenation
- Always work from the raw paste. The fixer needs original line endings. If you already joined the lines, the split-word evidence is gone.
- Remove soft hyphens separately. Invisible soft hyphens need the invisible character detector; this tool handles only visible line-end hyphens.
- Trust the word count. The rejoined-words count tells you the scale of the fix. Zero rejoins on hyphen-heavy text means the paste may have lost its line breaks already.
- Review proper names. When you enable rejoining before capitals, scan the output for names and sentence starts that should have kept their hyphens. Two minutes of review beats publishing a mangled name.
- Chain the tools in order. Soft hyphens out, visible hyphens rejoined, line breaks fixed, spacing collapsed. Each step prepares the input for the next.
Frequently asked questions
Will this tool remove hyphens that belong in the word?
No. It only removes hyphens at line ends, which are the ones PDFs add when wrapping text. Real hyphens inside a line, such as in well-known or state-of-the-art, are never touched. Hyphens before capitalized words are also preserved by default, protecting proper names.
How does it know which hyphens to remove?
It only acts on a hyphen at the very end of a line followed by a lowercase letter starting the next line, the classic signature of PDF line-break hyphenation. When in doubt, the tool errs on the side of keeping the hyphen.
What if the next line starts with a capital letter?
By default the tool keeps the hyphen, because a capital often signals a proper name or a real compound. If your document has many split words starting with capitals, tick also rejoin before capitals to join those too, then review the output.
Does it handle soft hyphens too?
Soft hyphens are invisible Unicode characters, a different problem. Use the invisible character detector to find and remove them, then run this tool for the visible line-end hyphens that remain.
Is my text uploaded anywhere?
No. All processing happens in your browser with JavaScript, so your documents stay private on your own device. You can safely fix confidential reports, legal filings, and unpublished manuscripts here.
Ready to try it yourself? It's free, no signup required.
Try the free PDF Hyphen Fixer →