Easy way to get data between tags of xml or html files in python?

Question

I am using Python and need to find and retrieve all character data between tags:

<tag>I need this stuff</tag>

I then want to output the found data to another file. I am just looking for a very easy and efficient way to do this.

If you can post a quick code snippet to portray the ease of use. Because I am having a bit of trouble understanding the parsers.

ghostdog74 · Accepted Answer · 2010-01-20 00:00:54Z

8

without external modules, eg

>>> myhtml = """ <tag>I need this stuff</tag>
... blah blah
... <tag>I need this stuff too
... </tag>
... blah blah """
>>> for item in myhtml.split("</tag>"):
...   if "<tag>" in item:
...       print item [ item.find("<tag>")+len("<tag>") : ]
...
I need this stuff
I need this stuff too

answered Jan 20, 2010 at 0:00

ghostdog74

346k62 gold badges264 silver badges349 bronze badges

Sign up to request clarification or add additional context in comments.

1 Comment

ghostdog74 Over a year ago

the so called "best" method is subjective. Depends on what is to be solved. If the task is simple, then use simple methods

Andrew Hare · Accepted Answer · 2010-01-19 23:10:44Z

Beautiful Soup is a wonderful HTML/XML parser for Python:

Beautiful Soup is a Python HTML/XML parser designed for quick turnaround projects like screen-scraping. Three features make it powerful:

Beautiful Soup won't choke if you give it bad markup. It yields a parse tree that makes approximately as much sense as your original document. This is usually good enough to collect the data you need and run away.

Beautiful Soup provides a few simple methods and Pythonic idioms for navigating, searching, and modifying a parse tree: a toolkit for dissecting a document and extracting what you need. You don't have to create a custom parser for each application.

Beautiful Soup automatically converts incoming documents to Unicode and outgoing documents to UTF-8. You don't have to think about encodings, unless the document doesn't specify an encoding and Beautiful Soup can't autodetect one. Then you just have to specify the original encoding.

Aiden Bell · Accepted Answer · 2010-01-19 23:11:59Z

2

I quite like parsing into element tree and then using element.text and element.tail.

It also has xpath like searching

>>> from xml.etree.ElementTree import ElementTree
>>> tree = ElementTree()
>>> tree.parse("index.xhtml")
<Element html at b7d3f1ec>
>>> p = tree.find("body/p")     # Finds first occurrence of tag p in body
>>> p
<Element p at 8416e0c>
>>> p.text
"Some text in the Paragraph"
>>> links = p.getiterator("a")  # Returns list of all links
>>> links
[<Element a at b7d4f9ec>, <Element a at b7d4fb0c>]
>>> for i in links:             # Iterates through all found links
...     i.attrib["target"] = "blank"
>>> tree.write("output.xhtml")

answered Jan 19, 2010 at 23:11

Aiden Bell

28.4k4 gold badges77 silver badges121 bronze badges

Comments

Shravya K · Accepted Answer · 2017-08-16 09:45:37Z

1

This is how I am doing it:

    (myhtml.split('<tag>')[1]).split('</tag>')[0]

Tell me if it worked!

answered Aug 16, 2017 at 9:45

Shravya K

371 silver badge3 bronze badges

Comments

torger · Accepted Answer · 2010-01-20 06:20:16Z

0

Use xpath and lxml;

from lxml import etree

pageInMemory = open("pageToParse.html", "r")

parsedPage = etree.HTML(pageInMemory)

yourListOfText = parsedPage.xpath("//tag//text()")

saveFile = open("savedFile", "w")
saveFile.writelines(yourListOfText)

pageInMemory.close()
saveFile.close()

Faster than Beautiful soup.

If you want to test out your Xpath's - I find FireFox's Xpather extremely helpful.

Further Notes:

edited Jan 20, 2010 at 6:20

answered Jan 20, 2010 at 6:15

torger

2,3384 gold badges28 silver badges36 bronze badges

1 Comment

Mallik Over a year ago

i want to remove the <italic>,</italic>,<bold>or</bold> tags for the text with in the <abstract> </abstract> tags. so how to achieve it. for example <abstract xml:lang="en"> <p> Abstract-As <italic>a</italic> promising method dealing with various wireless channel performance fading, <bold>MIMO</bold> technology is effective in supporting reliable, high-data-rate transmission. While the wireless system must cope with the space-time correlation and fast Rician fading.</abstract> Thanks in advance

E.G. Cortes · Accepted Answer · 2017-05-16 23:11:19Z

0

def value_tag(s):
    i = s.index('>')
    s = s[i+1:]
    i = s.index('<')
    s = s[:i]
    return s

edited May 16, 2017 at 23:11

answered May 16, 2017 at 23:05

E.G. Cortes

7512 bronze badges

Collectives™ on Stack Overflow

Easy way to get data between tags of xml or html files in python?

6 Answers 6

1 Comment

Comments

Comments

Comments

1 Comment

Comments

Your Answer

Linked

Hot Network Questions

Collectives™ on Stack Overflow

6 Answers 6

1 Comment

Comments

Comments

Comments

1 Comment

Comments

Your Answer

Sign up or log in

Post as a guest

Linked

Related