Strip XML Tags – Extract Text from XML

Extract the text content from XML, RSS feeds, SVG or config files, without the markup.

Loading tool…

About the Strip XML Tags

The input is parsed as XML with DOMParser, then the text of every element is collected in document order. CDATA sections are read as text, so <![CDATA[<b>bold</b>]]> yields <b>bold</b> literally. Entities like &amp; and numeric references like &#169; are decoded. Comments, processing instructions and the XML declaration are dropped, and attribute values are not included.

With One line per element, each element's text goes on its own line, which turns records into readable lists, for example the titles, descriptions and dates of an RSS feed. Turn it off to join everything with single spaces. Collapse extra whitespace removes the indentation used for pretty-printing.

If the document is not well-formed (a missing closing tag, a stray &), the tool falls back to removing anything that looks like a tag and says so in the status line, so you still get text from broken fragments. Nothing leaves your browser.

How to use it

  1. Paste XML or open a file.
  2. Choose one line per element or a single paragraph.
  3. Copy or download the text.

Frequently asked questions

Are attribute values kept?
No, only element text. Convert with XML to JSON or XML to CSV if you need attributes.
What happens to CDATA?
Its content is kept as plain text, without the CDATA markers. Markup inside CDATA is kept literally, not stripped.
Does it work on SVG or XHTML?
Yes, both are XML. For HTML that is not well-formed XML, use Strip HTML Tags instead.

Related tools