TinyFileLab

Home›Developer›Remove duplicate lines

Runs in the page — your list never leaves it

Remove duplicate lines

Paste the list. The repeats disappear as you type, and the order you had is left alone unless you ask otherwise.

Your lines 0 lines
Result 0 lines
What to keep
Nothing to do yet
Nothing is transmittedThe page does the work locally. No request carries your data anywhere.
No length limitPaste as much as you like — the only ceiling is your own memory.
No account, no ads in the wayOpen the page, use it, close it. Nothing to sign up for.

“Remove duplicates” means four different things

Most tools guess which one you meant and give you no way to say otherwise. This one asks. Keep the first appearance is the ordinary case: every distinct line survives, in the order it first showed up, and later repeats are dropped. Keep the last does the same but favours the final occurrence, which is what you want when a log or an export appends corrections after the original row.

The other two are filters rather than de-duplication. Only lines that appear exactly once throws away anything that was repeated at all — useful for finding the entries in a list that are not shared with anything else. Only the lines that were repeated does the reverse and shows you what the duplicates actually were, which is how you check a mailing list or an ID export before deciding what to do about it.

Order is preserved unless you ask

Sorting is the classic shortcut for this job — sort -u on the command line, or Excel's Remove Duplicates followed by a sort — and it silently destroys any ordering that carried meaning. Log lines in time order, a ranked list, a CSV whose rows correspond to something else: all scrambled to make the de-duplication easier to implement.

Here the original order is kept by default. The sort option exists because sometimes you do want it, and when it is on it sorts naturally, so item2 comes before item10 rather than after it. That is almost always what a human means by sorted and almost never what a plain string comparison gives you.

The two settings that change the answer

Ignore leading and trailing spaces is on by default and it matters more than it sounds. A line copied out of a spreadsheet column or an indented block of code often carries invisible trailing whitespace, and alice@example.com followed by a single space is a different string from alice@example.com as far as any exact comparison is concerned. With this on, they count as the same line; the line that survives keeps its original spacing.

Ignore capitals is off by default because it is not always safe. Email addresses and domain names are case-insensitive in practice, so turning it on is right for those. Usernames, passwords, file paths on Linux, base64 strings and hashes are not, and folding case there merges two genuinely different values. The counter under the buttons tells you how many distinct values were found, which is the quickest way to see whether a setting changed the answer.

Working on a file rather than a paste

You can drag a text file straight onto the left-hand box and it is read into place. Nothing is uploaded when you do — the file is read by the page through the browser's own file API, in the same way it would read a file you picked from a dialog.

There is no length limit beyond your own memory, and comparison uses a hash map rather than repeated scanning, so a hundred thousand lines de-duplicate as fast as a hundred. Empty lines are skipped by default and counted separately in the summary, since a list of blank rows collapsing into one is rarely what anyone wants.

Common questions

Does it change the order of my lines?

No, not unless you tick Sort the result. The de-duplicated list comes back in the order you gave it, which is what separates this from sorting the list and taking the unique values.

What counts as a duplicate?

Two lines whose text matches after the options you have set are applied. By default leading and trailing spaces are ignored and capitals are respected, so \u201cAlice\u201d and \u201calice\u201d are different until you tick Ignore capitals.

Can I see which lines were the duplicates?

Yes. Choose \u201cOnly the lines that were repeated\u201d and you get one copy of each line that appeared more than once, which is usually the first thing you want when auditing a list.

How is this different from Excel's Remove Duplicates?

Excel works on cells in a sheet and edits in place, which makes it awkward for a list you have pasted from somewhere else. This works on plain text, shows the result beside the input, and lets you keep the last occurrence instead of the first.

Is there a limit on how much I can paste?

None that the page imposes. Matching uses a hash map, so the work grows in step with the number of lines rather than exploding. Hundreds of thousands of lines are fine on an ordinary machine.

Is my list sent anywhere?

No. Everything runs in this page, and there is no request carrying your text at any point. That matters when the list is customer emails or internal identifiers — disconnect from the internet and the page keeps working.