Chapter 2. Advanced HTML Parsing

When Michelangelo was asked how he could sculpt a work of art as masterful as his David, he is famously reported to have said: âIt is easy. You just chip away the stone that doesnât look like David.â

Although web scraping is unlike marble sculpting in most other respects, we must take a similar attitude when it comes to extracting the information weâre seeking from complicated web pages. There are many techniques to chip away the content that doesnât look like the content that weâre searching for, until we arrive at the information weâre seeking. In this chapter, weâll take look at parsing complicated HTML pages in order to extract only the information weâre looking for.

You Donât Always Need a Hammer

It can be tempting, when faced with a Gordian Knot of tags, to dive right in and use multiline statements to try to extract your information. However, keep in mind that layering the techniques used in this section with reckless abandon can lead to code that is difficult to debug, fragile, or both. Before getting started, letâs take a look at some of the ways you can avoid altogether the need for advanced HTML parsing!

Letâs say you have some target content. Maybe itâs a name, statistic, or block of text. Maybe itâs buried 20 tags deep in an HTML mush with no helpful tags or HTML attributes to be found. Letâs say you dive right in and write something like the following line to attempt extraction:

bsObj.findAll("table" ...

Get Web Scraping with Python now with the O’Reilly learning platform.

O’Reilly members experience books, live events, courses curated by job role, and more from O’Reilly and nearly 200 top publishers.

Start your free trial

Web Scraping with Python by Ryan Mitchell

Chapter 2. Advanced HTML Parsing

You Donât Always Need a Hammer

Don’t leave empty-handed

It’s yours, free.

Check it out now on O’Reilly

Chapter 2. Advanced HTML Parsing

You Donât Always Need a Hammer

Don’t leave empty-handed

It’s yours, free.

Check it out now on O’Reilly

You Donât Always Need a Hammer