TinyFileLab

Home›Convert›HTML to Markdown

Parsed in the page — scripts never execute

Convert HTML to Markdown

Paste a page, an email body or an export, and get Markdown back. Structure is kept, presentation is thrown away, and the scripts never had a chance to run.

HTML in paste it, or drop an .html file

Paste HTML on the left. Scripts, styles and comments are dropped before anything else happens.

Markdown out
Nothing is transmittedThe page does the work locally. No request carries your data anywhere.
No length limitPaste as much as you like — the only ceiling is your own memory.
No account, no ads in the wayOpen the page, use it, close it. Nothing to sign up for.

Converting down, not across

Markdown expresses a fraction of what HTML can. That makes this conversion a deliberate act of discarding: a <div class="callout"> has no Markdown equivalent, so it becomes its contents, and the class goes. The same is true of every inline style, every data attribute, every wrapper the original page used for layout. What survives is the structure that Markdown has a word for — headings, paragraphs, emphasis, code, quotes, lists, tables, links and images.

That is normally exactly what you want. The usual reason to run this conversion is that you have HTML from somewhere — a CMS export, an email body, a scraped article, output from a rich-text editor — and you want it in a repository, a documentation site, or an issue tracker where the presentation will come from somewhere else.

The parts that are easy to get wrong

Three things separate a usable conversion from a mess. Nested lists need their indentation to match the marker width, or the nesting silently flattens when the Markdown is rendered again — this page indents by the width of the parent's marker, including the space, which is what every Markdown implementation expects. Tables need the header separator row and a consistent column count; rows that were short in the HTML get padded rather than producing a broken table. And text escaping matters more than it looks: a paragraph that happens to start with a hyphen or contain an asterisk will turn into a list item or italics on the round trip unless the character is escaped.

Underscores are deliberately left alone. Escaping every one of them is technically safer and makes any document containing snake_case identifiers unreadable, so the trade is made the other way.

Code blocks keep their language

A <pre><code class="language-python"> block becomes a fence tagged python. That is the convention emitted by every syntax highlighter and read by GitHub, GitLab and most static site generators, so highlighting survives the trip. The contents are taken verbatim — no escaping, no whitespace collapsing — and if the code itself contains a run of backticks the fence grows until it is longer than anything inside it.

Scripts never run

The HTML is parsed with DOMParser, which builds a document tree without executing anything: no scripts, no inline event handlers, no network requests for images or stylesheets. Script, style, iframe, SVG and template elements are then removed outright, along with every comment, before the walk begins. This is the safe way to handle HTML you did not write, and it is the reason a converter that runs on your own machine is preferable to one that fetches your paste to a server and renders it there.

What it will not do

Definition lists, footnote markup and complex tables with merged cells have no clean Markdown form; their text is preserved but the structure is not. Elements Markdown cannot express are unwrapped rather than dropped, so you keep the words. If a page relies on layout to carry meaning — a multi-column grid, a diagram made of positioned divs — no converter can rescue that, and this one does not pretend to.

For the opposite direction, the Markdown to HTML converter is the mirror of this page and uses the same conventions, so a round trip through both is stable for anything within Markdown's vocabulary.

Common questions

Does it handle nested lists?

Yes, to any depth. Each level is indented by the width of its parent's marker, which is what Markdown renderers expect — get that wrong and the nesting collapses when the file is rendered again.

What happens to tables?

They become GitHub-style pipe tables with a header separator row. Short rows are padded so the column count stays consistent. Merged cells cannot be represented and lose their structure.

Are my scripts executed when I paste HTML?

No. Parsing uses DOMParser, which never executes anything, and script, style, iframe and SVG elements are removed before conversion starts.

Why are some characters preceded by a backslash?

Because they mean something in Markdown. An asterisk or a leading hyphen in ordinary text would become emphasis or a list item on the way back, so they are escaped. Underscores are left alone to keep snake_case readable.

Do code blocks keep their language?

Yes, if the source used the standard class="language-x" convention. The fence is tagged with it, and grows longer than any run of backticks inside the code.

Can I convert a whole saved web page?

Yes — drop the .html file onto the input box. You will get the body content; navigation and footers convert too, so expect to delete some of it.

Is anything uploaded?

No. The parsing and conversion happen in this page, which also means there is no size limit beyond your own memory.