A lot of documents never pass through Word. Meeting notes and READMEs are written in Markdown, drafts come out of ChatGPT or Claude as Markdown, reports are exported as HTML pages, and logs are plain text. Sooner or later one of them has to be handed in as a PDF, and sometimes as a PDF that looks like it came off a scanner.
Look Scanned could already open .md, .html and .txt files, but only as beta formats. The whole file landed on one tall page. For anything longer, the advice was to print it to PDF from your browser first and scan that. That detour is gone. All three formats are now laid out as A4 pages inside your browser, and the beta label has come off.
Real pages, with the breaks where you put them
Whatever you open, the result is a stack of pages rather than one long strip:
- Text flows onto the next page on its own.
- Page breaks you added in the file are kept.
- A long table carries on to the next page, and a row is never cut in half.
- A long code block continues on the next page, and long lines wrap instead of running off the edge.
- An image taller than the page is shrunk to fit.
Markdown
Markdown files get the formatting you know from GitHub: headings, lists, tables, quotes and fenced code blocks. Images show up when they are embedded in the file or linked with a full https:// address. An image referenced by a relative path such as images/chart.png can't be found, because the file is opened on its own, without its folder.
Markdown has no syntax of its own for a page break, so you drop in a line of HTML wherever the next page should start:
<div style="break-before: page"></div>
The older page-break-before: always works as well.
If an AI assistant wrapped its whole answer in a code block, remove the outer triple backticks before saving the file, or the whole document comes out as one code block.
HTML
An HTML file is laid out with the styles written inside it: its <style> blocks and style attributes. Fonts, colours, spacing and your break-before or page-break-before rules all apply. A stylesheet kept in a separate file is not loaded. Images and web fonts with a full https:// address are fetched from wherever they live, so they need an internet connection. A relative path such as page_files/photo.jpg can't be found, for the same reason as in Markdown.
The page is treated as a document to print, not a web page to run. Scripts and embedded frames are removed, and nothing on the page can run or be submitted. Web pages saved by your browser as .mhtml or .mht open too, with the pictures and styles saved inside the archive. Nothing is fetched from the web for those.
Plain text
A .txt, .text or .log file keeps its line breaks and indentation. Lines wider than the page wrap instead of being clipped. Files are read as UTF-8, and any tags or code in the text appear exactly as typed.
How to turn one into a scanned PDF
- Open Markdown to Scanned PDF, HTML to Scanned PDF or TXT to Scanned PDF. Any of the three pages opens any of the three formats.
- Select your file and wait a moment while the pages are laid out.
- Flip through the preview. Check where the pages break and that tables and code blocks look right.
- Adjust the scan effect. For small text or code, keep blur and noise low so the text stays easy to read.
- Click "Generate Scanned PDF", then "Download Scanned PDF".
You can also save each page as a JPG or PNG. Our guide to PDF, JPG and PNG explains when each one makes sense. With Pro, Bulk Scan takes a batch of files at once.
Your file stays on your device
The file is laid out and scanned on your own device. It isn't uploaded to a server. The one exception is an HTML or Markdown file that links to pictures or fonts on the web: those are fetched from the sites that host them, just as they would be if you opened the page normally.
Good to know
- Pages are always A4 with one-inch margins. A paper size or margin set in the file's own CSS is ignored. If you need US Letter, choose it under Paper Settings and each page is fitted onto it.
- Fonts come from your device. Unless the HTML's own styles load web fonts, the text is set in the fonts installed on your computer or phone, so the same file can break pages differently on another machine.
- Formulas and diagrams stay as text. Maths between dollar signs and diagram code such as Mermaid appear as the text you typed. If you need them drawn, export a PDF from an editor that renders them and scan that PDF instead.
- Very long files can run out of time. Layout has to finish within about 30 seconds, so a huge log file or a book-length Markdown file may stop with an error. Split it into smaller files first.
- The output is made of images. Like a real scan, the text can't be selected or searched. Keep the original file for editing.
Had a Markdown or HTML file come out as one long page before? Try it again.