How to web scrap on Python with BeautifulSoup?

Author :

React :

Comment

To make web scraping in Python with BeautifulSoup, you need two libraries: requests to download the web page, and beautifulsoup4 to extract data from the HTML code.

Install them using the command pip install requests beautifulsoup4, then follow the step-by-step guide below to scrape a website.

Python web scraping script using BeautifulSoup to extract data from a web page
Web scraping on Python with BeautifulSoup. ©Christina for Alucare.fr

Prerequisites for scraping Python with BeautifulSoup

Before you get started, you don't need to be an expert. You just need to know how to read and follow a script Python That's more than enough to get started.

Here's what you need to install to get started:

  • Install Python as well as a development environment.
  • Install pip, the tool that lets you add Python libraries with a single command.
  • Install BeautifulSoup :
pip install beautifulsoup4

Pay attention to the package name: on PyPI, that's right beautifulsoup4 that you need to write. The module you'll import into your code is called bs4.

  • Install Requests To download the web pages:
pip install requests
  • Install lxml (optional), a faster parser that BeautifulSoup uses to parse HTML:
pip install lxml

How to web scrap with Python and BeautifulSoup?

Here is a complete project for recovering the title of a web page and all the links that it contains.

Python web scraping script using the BeautifulSoup library to extract HTML data
Web Scraping in Python with BeautifulSoup. ©Christina for Alucare.fr

Step 1: Retrieve page content with Requests

To download a web page, you send a HTTP GET request to a URL. That's the library's role Requests.

For each request, the server returns a status code which tells you whether everything went well. The main ones you should know:

  • 200 : success.
  • 301 / 302 redirection.
  • 404 page not found.
  • 500 internal server error.

With Requests, you can verify this result using the attribute .status_code. Here's a concrete example: this code sends a request to a website, checks the status code, and then displays a snippet of the HTML content if everything goes well.

import requests

# Target URL
url = "https://bonjour.com"

# Send a GET request
response = requests.get(url)

# Check status code
if response.status_code == 200:
    print("Success: the page has been retrieved!")
    html = response.text # HTML content of the page
    print("Extract HTML content:")
    print(html[:500]) # displays only the first 500 characters
else:
    print(f "Error: status code {response.status_code}")

Step 2: Analyze HTML code with BeautifulSoup

When you retrieve the content of a page using response.text, you get a simple string: the entire HTML code from the page, but it's unusable as is. To work with it easily, create an object BeautifulSoup, which converts this raw HTML into a structure that you can browse and analyze.

Always indicates a parser, for example "html.parser. BeautifulSoup then parses the HTML correctly, without displaying any warnings. You can also use lxml, which is faster and must be installed separately.

from bs4 import BeautifulSoup
import requests

url = "https://bonjour.com"
response = requests.get(url)
html = response.text

# Specifying the parser is recommended
soup = BeautifulSoup(html, "html.parser")

Step 3: Find and extract elements

The HTML is now a BeautifulSoup object. You can search for and retrieve the data that interest you, tag by tag. Three methods cover most needs.

Using find() and find_all()

find() returns the first element found. find_all() returns the full list the corresponding elements.

# Recover title <h1>
h1 = soup.find("h1")
print(h1.get_text())

# Retrieve all links <a>
liens = soup.find_all("a")
for lien in liens:
    print(lien.get_text(), lien.get("href"))

Target elements by attribute

You can refine your search based on an HTML attribute, such as class, id or any other. To note : In Python, you write class_ and no class, to avoid a conflict with the reserved word in the language.

# Retrieve a div with a specific ID
container = soup.find("div",)

# Retrieve all links with a specific class
nav_links = soup.find_all("a", class_="nav-link")

Using CSS Selectors with select()

For more specific searches, the method select() accepts CSS selectors. You can target specific parts of a page without having to go through all the HTML manually.

# All links in article titles
links_articles = soup.select("article h2 a")

# All <a> whose href attribute begins with "http".
links_http = soup.select('a[href^="http"]')

How to extract data from an HTML table with BeautifulSoup?

Extracting Data from an HTML Table Using BeautifulSoup in Python
Extract data from an HTML table with BeautifulSoup. ©Christina for Alucare.fr

Real-world use cases are often more complex than simply retrieving a title or a link. You'll need to handle the’structured data extraction Like tables and lists, the page numbers, and common scraping errors.

Extract tables and lists

Websites often present their data in HTML tables (<table>, <tr>, <th>, <td>) or lists (

    /
      with
    1. ). To convert these structures into usable data, you need to iterate through them line by line or element by element.

      To extract a HTML tablethe principle is simple:

      • Retrieve the headers (<th>) to identify column headings.
      • Go through each line (<tr>) and search for cells (<td>) that contain the data.
      • Store the information in a list or a dictionary.

      For a HTML list (

        Where
          ) :

          • Find all the tags
          • with find_all.
          • Retrieve their content (text or link) and add it to a Python list.

          Here's an example with a table:

          html = """
          <table>
            <tr>
              <th>Last name</th>
              <th>Age</th>
              <th>Town</th>
            </tr>
            <tr>
              <td>Alice</td>
              <td>25</td>
              <td>Paris</td>
            </tr>
            <tr>
              <td>Bob</td>
              <td>30</td>
              <td>Lyon</td>
            </tr>
          </table>
          """
          
          # Create BeautifulSoup object
          soup = BeautifulSoup(html, "html.parser")
          
          # Extract headers from array
          headers = [th.get_text(strip=True) for th in soup.find_all("th")]
          print("Headers:", headers)
          
          # Extract data rows (skip 1st row as these are the headers)
          rows = []
          for tr in soup.find_all("tr")[1:]:
              cells = [td.get_text(strip=True) for td in tr.find_all("td")]
              if cells:
                  rows.append(cells)
          
          print("Lines :", rows)
          

          Here, find_all("th") retrieves the headings and find_all("td") retrieves the cells from each row. You loop through the <tr> to rebuild the table row by row.

          Here's an example on a list:

          from bs4 import BeautifulSoup
          
          html_list = """
          
          • Apple
          • Banana
          • Orange
          """ soup = BeautifulSoup(html_list, "html.parser") # Retrieve list items items = [li.get_text(strip=True) for li in soup.find_all("li")] print("Extracted list:", items) # ["Apple", "Banana", "Orange"]

          Each

        1. is directly converted into Python list element, which gives the result ["Apple", "Banana", "Orange"].

          Manage pagination and links

          Often, the data doesn't fit on a single page. It is distributed via links to “next page” or a numbered pagination (?page=1, ?page=2, etc.). In both cases, you must curl to retrieve all the pages and merge the data.

          Example using a page parameter:

          import time
          import requests
          from bs4 import BeautifulSoup
          
          # Example of URL with pagination
          BASE_URL = "https://bonjour.com/articles?page={}"
          HEADERS = {"User-Agent": "Mozilla/5.0"}
          
          all_articles = []
          
          # Assume 5 pages to browse
          for page in range(1, 6):
              url = BASE_URL.format(page)
              r = requests.get(url, headers=HEADERS, timeout=20)
              if r.status_code == 200:
                  soup = BeautifulSoup(r.text, "html.parser")
                  # Extract article titles
                  articles = [h2.get_text(strip=True) for h2 in soup.find_all("h2", class_="title")]
                  all_articles.extend(articles)
              else:
                  print(f "Error on page {page} (code : {r.status_code})")
              time.sleep(1.0) # politeness
          
          print("Articles retrieved :", all_articles)

          A brief step-by-step explanation:

          • Prepare the’URLs with a spot {} to insert the page number.
          BASE_URL = "https://bonjour.com/articles?page={}
          • Some websites block requests that do not include a browser ID. Add a User-Agent Avoid being mistaken for a bot.
          headers = {"User-Agent": "Mozilla/5.0"}
          requests.get(url, headers=headers) 
          • Page Footer 1 to 5.
          for page in range(1, 6):
          • Retrieve the HTML from the page using requests.
          requests.get(url)
          • Limits the wait time if the site does not respond (20-second timeout).
          requests.get(url, timeout=20)
          • Analyze the page with BeautifulSoup.
          BeautifulSoup(response.text, "html.parser")
          • Retrieves all the article titles.
          find_all("h2", class_="title")
          • Adds the items found to a master list.
          all_articles.extend(articles)
          • Enter a break between each request to avoid overloading the server and getting banned.
          time.sleep(1.0)

          After the loop, all_articles contains all 5 page titles.

          Common mistakes and challenges

          Web scraping isn't just about pressing a button. You'll run into some common obstacles:

          • HTTP errors : 404 (page not found), 403 (no entry) and 500 (server-side error).

          Management example:

          response = requests.get(url)
          if response.status_code == 200:
              print("Page retrieved successfully")
          elif response.status_code == 404:
              print("Error: Page not found")
          else:
              print("Status code returned:", response.status_code)
          
          • Sites that block scraping : Some detect automated requests and block access.
          • Dynamic pages (JavaScript) : BeautifulSoup only reads the Static HTML. If the page loads its content using JavaScript, you won't see anything. In that case, you should use a tool like Selenium Where Playwright.

          To scrape effectively without getting blocked or damaging the site, here are the best practices:

          • Respect the robots.txt file of the website you want to analyze.
          • Set up some deadlines between queries with time.sleep() so as not to overload the server.
          • Uses proxies and spin them.
          • Change your User-Agent.

          Web scraping with Selenium and BeautifulSoup?

          Web scraping with Selenium and BeautifulSoup on Chrome using Python.
          Web scraping with Selenium and BeautifulSoup on Chrome. ©Christina for Alucare.fr

          Selenium control a real browser, runs the JavaScript and displays the page as if a human were browsing it. BeautifulSoup Then analyze the HTML once the page has fully loaded, and extract all the data you want.

          Step 1: Install Selenium and BeautifulSoup

          Here, instead of requests, you use Selenium to retrieve the page's content. Install the two libraries using this command:

          pip install selenium beautifulsoup4

          Next, you need to download a WebDriver tailored to your browser version. For Google Chrome, it's ChromeDriver. You have two options: place it in the same folder as your Python script, or’add to the PATH environment variable of your system.

          Step 2: Configure Selenium

          Start by importing webdriver using Selenium to control the browser.

          from selenium import webdriver
          from selenium.webdriver.common.by import By

          Next, launch a browser. The browser opens the page and runs the JavaScript, in this case Chrome.

          driver = webdriver.Chrome()

          Tells the browser which URLs visit.

          driver.get("https://www.exemple.com")

          If the page takes a while to display certain elements, you can tell Selenium to wait up to 10 seconds.

          driver.implicitly_wait(10)

          Step 3: Retrieving page content

          Once the page has loaded, you retrieve the Full DOM : the HTML source code generated after the JavaScript is executed.

          html_content = driver.page_source

          Step 4: HTML analysis with BeautifulSoup

          Now pass this source code to BeautifulSoup To use it:

          from bs4 import BeautifulSoup
          
          # Create a BeautifulSoup object
          soup = BeautifulSoup(html_content, 'html.parser')
          
          # Example: Retrieve all headings from the page
          headings = soup.find_all('h2')
          for heading in headings:
              print(heading.get_text())

          BeautifulSoup offers powerful methods such as find(), find_all() and CSS selectors for target and extract specific HTML elements.

          Step 5: Closing the browser

          Very important: Always close your browser after running the program to free up resources.

          driver.quit()

          This way, you combine the power of Selenium to simulate human navigation (clicks, scrolls) combined with BeautifulSoup's efficiency in parsing HTML. This is the ideal method for scraping a site that loads its data via JavaScript.

          FAQs

          What's the best tool for web scraping in Python?

          There is no’scraping tool There's no one-size-fits-all solution—rather, solutions tailored to your project. Here are the three most commonly used ones:

          • BeautifulSoup : Simple and effective for parsing HTML and quickly extracting content. Ideal if you're just starting out or working on small projects.
          • Scrapy : a comprehensive framework designed to manage large data volumes with advanced features.
          • Playwright : Perfect for complex JavaScript-driven websites. It simulates a real browser and interacts with the page just like a human would.

          How do I use BeautifulSoup to extract the content of a div tag?

          With BeautifulSoup, you target a specific tag using a CSS selector. To extract the content of a tag <div>, follow these steps.

          1. Retrieve the web page using requests, then parse the HTML with BeautifulSoup.

          from bs4 import BeautifulSoup
          import requests
          
          url = "URL_DE_TON_SITE"  # Remplace par l'URL réelle
          reponse = requests.get(url)
          html_content = reponse.text
          
          soup = BeautifulSoup(html_content, "html.parser")

          2. Use the method select() using your CSS selector to target the tag <div>.

          To retrieve the first element, use soup.select_one. To retrieve all the elements, use soup.select. Here is an example of HTML code:

          <div>
            <h2>Titre de l'article</h2>
            <p>Voici le contenu du paragraphe.</p>
          </div>

          And here's an example using the CSS selector div.article :

          # Récupérer le premier div avec la classe "article"
          div_article = soup.select_one("div.article")
          
          # Afficher son contenu texte
          if div_article:
              print(div_article.get_text(strip=True))

          3. Remove the items inside the <div>.

          # Récupérer le titre à l'intérieur du div
          titre = soup.select_one("div.article h2").get_text()
          
          # Récupérer le paragraphe à l'intérieur du div
          paragraphe = soup.select_one("div.article p").get_text()
          
          print("Titre :", titre)
          print("Paragraphe :", paragraphe)

          How do I use Requests and BeautifulSoup together?

          These two libraries are additional : One downloads the page, the other analyzes it.

          1. requests sends an HTTP request and downloads the raw HTML code of the page.

          import requests
          
          url = "https://sitecible.com"
          response = requests.get(url)  # requête HTTP
          print(response.text)  # affiche le HTML brut

          At this point, you just have a huge block of text full of tags (<html>, <div>, <p>etc.).

          2. BeautifulSoup analyzes this raw HTML and converts it into a organized structure. You can then browse the page, locate the tags, and retrieve the data.

          from bs4 import BeautifulSoup
          
          soup = BeautifulSoup(response.text, "html.parser")  # analyse le HTML
          titre = soup.find("h1").get_text()  # extrait le contenu d'un <h1>
          print(titre)

          Why doesn't my web scraping code work on some sites?

          Sometimes your script won't retrieve anything. Some sites don't provide all their content directly in HTML: they use JavaScript to load the data dynamically. However, BeautifulSoup cannot parse the data returned by JavaScript. In that case, you should turn to tools like Playwright Where Selenium.

          What role does BeautifulSoup play in web scraping?

          BeautifulSoup plays the role of’HTML parser. It converts a page's code into a structured object that you can easily navigate. In short, it's the a converter between raw HTML and your Python code.

          Web Scraping: BeautifulSoup or Scrapy?

          BeautifulSoup and Scrapy are really different, even though both are used for web scraping.

          BeautifulSoup Scrapy
          A simple library for parsing HTML and extracting data. A comprehensive framework that handles the entire scraping process: requests, link tracking, pagination, data export, and error handling.

          BeautifulSoup is the perfect choice for beginners: it makes the’HTML data extraction in Python Quick and easy.

          If you'd rather keep the coding to a minimum, the tool Bright Data offers ready-to-use datasets to reduce the amount of coding required.

          👍Your opinion
          The article is informative
          The article is objective
          The article answers my question
          Content up to date
          🔍 Found any errors? Tell us where!

Found this helpful? Share it with a friend!

This content is originally in French (See the editor just below.). It has been translated and proofread in various languages using Deepl and/or the Google Translate API to offer help in as many countries as possible. This translation costs us several thousand euros a month. If it's not 100% perfect, please leave a comment for us to fix. If you're interested in proofreading and improving the quality of translated articles, don't hesitate to send us an e-mail via the contact form!
We appreciate your feedback to improve our content. If you would like to suggest improvements, please use our contact form or leave a comment below. Your feedback always help us to improve the quality of our website Alucare.fr


Alucare is an free independent media. Support us by adding us to your Google News favorites:

Post a comment on the discussion forum