No PDF text handy? Load a sample.

Samples contain real no-break spaces, hyphenated line ends, and emoji. Loading one replaces the text in the left box.

Text processing happens in your browser. Nothing you paste is uploaded, saved, or added to the page address. Results use LF line breaks, and the TXT download is UTF-8 without a byte order mark.

Cleaning rules
Join separator Space suits English. No space can suit text written without spaces between words. The tool does not detect the language.

Live result

What changed

Paste text to see what changes.

Prose boundaries joined
0
No-break spaces converted
0
Hyphen candidates pending
0
Hyphen candidates confirmed
0

Needs your decision

Line-end hyphenation

Paste text to look for words split by a hyphen at the end of a line.

    Check the result

    Review changes

    PDF copy and paste guide

    How this differs from a plain line-break remover

    A generic line-break remover turns every line break into a space. That is the wrong move for a bullet list, a code sample, or a word that a PDF split with a hyphen at the end of a line. This tool first sorts the pasted lines into ordinary prose, protected blocks, and blank lines, then cleans only the prose. For a plain whole-text replacement, use Remove Line Breaks. For the reasons PDF text pastes this way, read How to Clean Text Copied from a PDF.

    Everything is calculated again from your original paste each time you type, change a rule, or decide a hyphen. Rules never pile up on a previous result, and the left box always keeps your original text.

    What each rule does

    Line breaks are normalized first: CRLF and CR become LF, a line separator (U+2028) becomes one LF, and a paragraph separator (U+2029) becomes two. After that the rules below apply. A blank line is a line that holds only ordinary spaces, Tabs, or no-break spaces.

    RuleDefaultExact behavior
    Preserve detected listsOnA line with up to three ordinary spaces, then -, *, •, or digits followed by . or ), then a space, Tab, or no-break space starts a list. Every line from there to the next blank line is kept as pasted.
    Preserve indented textOnA non-blank line that starts with a Tab (even after a few spaces) or with four ordinary spaces is kept as pasted. Neighboring protected lines, and blank lines between them, form one block.
    Convert non-breaking spacesOnU+00A0, U+202F, and U+2007 in prose become U+0020. Other Unicode spaces are left alone.
    Clean ordinary spacingOnTrims U+0020 at both ends of each prose line and collapses two or more U+0020 into one. Tabs are not collapsed.
    Join wrapped prose linesOnMerges consecutive prose lines within a paragraph. A blank line, a protected block, or an undecided hyphen boundary stops the join.
    Join separatorSpaceThe text placed between joined lines: one ordinary space, or nothing.
    Normalize paragraph gapsOnReduces runs of blank lines between prose to one, and removes blank lines at the start and end. Blank lines inside a protected block are unchanged.

    Worked example

    This text has a doubled-space heading, two words split by a hyphen, a bullet list, and an indented command. With the default rules:

    Input Quarterly report summary The report was com- pleted on time, and the co- operation of every team was noted. Highlights: - Revenue grew 12% across all regions - Costs stayed flat npm test Default result Quarterly report summary The report was com- pleted on time, and the co- operation of every team was noted. Highlights: - Revenue grew 12% across all regions - Costs stayed flat npm test

    The heading and last prose line lost their repeated spaces. The list, its continuation line, and the indented command were not touched. The two hyphenated line ends stay exactly as copied because they are waiting for a decision. After you choose Join, remove hyphen for the first and Join, keep hyphen for the second, the paragraph becomes:

    The report was completed on time, and the co-operation of every team was noted.

    Why every line-end hyphen needs your decision

    A hyphen at the end of a line can be a typesetting split, as in docu- then ment, or a real hyphen, as in co- then operate. The characters are identical, so no tool can tell from the text alone. Guessing wrong either invents a word or deletes a real hyphen.

    The tool lists each place where a prose line ends with a letter and a hyphen and the next line starts with a lowercase letter. For each one you can keep it as copied (the default), join it and keep the hyphen, or join it and remove the hyphen. Each choice is undoable, and it applies only to that line break, never to every hyphen in the document. Undecided boundaries are still exported, exactly as copied, so the result is never silently guessed.

    Only ASCII letters are checked. Hyphenation in other languages, capitalized continuations, and lines inside protected blocks are not offered as candidates, so an empty list does not prove the text has no split words.

    What this tool cannot recognize

    • Paragraphs without blank lines. Many PDF viewers copy a whole page with no blank line between paragraphs, and some drop the indentation of code. Prose lines with no gap are then joined as one paragraph, a heading can be merged into the text after it, and a list marker protects every following line up to the next blank line, so lines after a numbered list stay unjoined. Look at the changes view before relying on the result, and switch off Preserve detected lists if the rest of the text should be joined.
    • Headings, captions, page numbers, headers, and footers. They are ordinary lines. They are not detected or removed, and a heading with no blank line after it can be merged into the next paragraph.
    • Multi-column pages and tables. Text copied across columns arrives in whatever order the PDF stores it. The tool cannot restore the reading order or the table cells.
    • Lists without markers and unusual indentation. Only the marker and indentation rules above are used. A paragraph that starts with a number and a period, such as a year, can be treated as a list, which keeps its lines separate rather than joining them.
    • Scanned PDFs and OCR mistakes. It only works on text that already exists. It does not read images, correct recognition errors, or restore characters the PDF never contained.

    Line breaks and invisible characters

    Browsers convert CRLF and CR to LF when text is pasted into a text box, so the original box and the result both use LF, and the downloaded TXT uses LF. The tool does not claim to keep the original line-ending bytes.

    Soft hyphens (U+00AD), zero-width spaces (U+200B), word joiners (U+2060), and U+FEFF are reported but not removed, because removing them can change how a word is searched or displayed. Joiners that hold emoji and some scripts together (U+200C, U+200D) and variation selectors are never touched. To inspect or clean the reported characters, use the zero-width character remover. To check no-break spacing in other text, use Remove Non-Breaking Spaces.

    How to clean text copied from a PDF

    1. Paste the text into the left box, or load a sample to see how the rules behave.
    2. Leave the default rules on, or switch off protection for lists or indented text if the tool kept lines you wanted joined.
    3. Go through the line-end hyphenation list and choose an action for each boundary you can judge. Leave the rest as copied.
    4. Open View changes to check joined lines, converted spaces, and protected blocks.
    5. Copy the result or download it as a TXT file. Copy original always returns your untouched paste.

    Frequently asked questions

    Does it upload my PDF or text?

    No file is uploaded, and this tool only takes pasted text. The cleaning runs in your browser, and the text is not saved or put in the page address. The site as a whole loads analytics and advertising scripts, described in the privacy policy.

    Why were some lines not joined?

    They were inside a protected list or indented block, they sat on either side of a blank line, or they ended with a hyphen that is waiting for your decision. Open View changes and tick the box to list protected blocks.

    Can it fix a PDF that copies every word separately or in the wrong order?

    No. It cleans spacing and line wrapping in text that was copied correctly. It cannot rebuild reading order, columns, or tables from plain text.

    What does Copy original do?

    It copies the text exactly as it was pasted, without running any cleaning. Resetting the rules recalculates the default result but does not change the original.

    Is there a size limit?

    Each paste can be up to 5 MiB of UTF-8 text and 100,000 lines. Larger input is refused with a message instead of being cut off.

    Related tools and guides

    Use Remove Line Breaks to join whole text without protecting anything, Remove Empty Lines for blank-line cleanup, Remove Non-Breaking Spaces for no-break spacing alone, and the PDF cleanup guide for manual steps and the limits of copying from PDFs.