To make web scraping in Python with BeautifulSoup, you need two libraries: requests to download the web page, and beautifulsoup4 to extract data from the HTML code.
Install them using the command pip install requests beautifulsoup4, then follow the step-by-step guide below to scrape a website.

Prerequisites for scraping Python with BeautifulSoup
Before you get started, you don't need to be an expert. You just need to know how to read and follow a script Python That's more than enough to get started.
Here's what you need to install to get started:
- Install Python as well as a development environment.
- Install pip, the tool that lets you add Python libraries with a single command.
- Install BeautifulSoup :
pip install beautifulsoup4
Pay attention to the package name: on PyPI, that's right beautifulsoup4 that you need to write. The module you'll import into your code is called bs4.
- Install Requests To download the web pages:
pip install requests
- Install lxml (optional), a faster parser that BeautifulSoup uses to parse HTML:
pip install lxml
How to web scrap with Python and BeautifulSoup?
Here is a complete project for recovering the title of a web page and all the links that it contains.

Step 1: Retrieve page content with Requests
To download a web page, you send a HTTP GET request to a URL. That's the library's role Requests.
For each request, the server returns a status code which tells you whether everything went well. The main ones you should know:
- 200 : success.
- 301 / 302 redirection.
- 404 page not found.
- 500 internal server error.
With Requests, you can verify this result using the attribute .status_code. Here's a concrete example: this code sends a request to a website, checks the status code, and then displays a snippet of the HTML content if everything goes well.
import requests
# Target URL
url = "https://bonjour.com"
# Send a GET request
response = requests.get(url)
# Check status code
if response.status_code == 200:
print("Success: the page has been retrieved!")
html = response.text # HTML content of the page
print("Extract HTML content:")
print(html[:500]) # displays only the first 500 characters
else:
print(f "Error: status code {response.status_code}")
Step 2: Analyze HTML code with BeautifulSoup
When you retrieve the content of a page using response.text, you get a simple string: the entire HTML code from the page, but it's unusable as is. To work with it easily, create an object BeautifulSoup, which converts this raw HTML into a structure that you can browse and analyze.
Always indicates a parser, for example "html.parser. BeautifulSoup then parses the HTML correctly, without displaying any warnings. You can also use lxml, which is faster and must be installed separately.
from bs4 import BeautifulSoup
import requests
url = "https://bonjour.com"
response = requests.get(url)
html = response.text
# Specifying the parser is recommended
soup = BeautifulSoup(html, "html.parser")
Step 3: Find and extract elements
The HTML is now a BeautifulSoup object. You can search for and retrieve the data that interest you, tag by tag. Three methods cover most needs.
Using find() and find_all()
find() returns the first element found. find_all() returns the full list the corresponding elements.
# Recover title <h1>
h1 = soup.find("h1")
print(h1.get_text())
# Retrieve all links <a>
liens = soup.find_all("a")
for lien in liens:
print(lien.get_text(), lien.get("href"))
Target elements by attribute
You can refine your search based on an HTML attribute, such as class, id or any other. To note : In Python, you write class_ and no class, to avoid a conflict with the reserved word in the language.
# Retrieve a div with a specific ID
container = soup.find("div",)
# Retrieve all links with a specific class
nav_links = soup.find_all("a", class_="nav-link")
Using CSS Selectors with select()
For more specific searches, the method select() accepts CSS selectors. You can target specific parts of a page without having to go through all the HTML manually.
# All links in article titles
links_articles = soup.select("article h2 a")
# All <a> whose href attribute begins with "http".
links_http = soup.select('a[href^="http"]')
How to extract data from an HTML table with BeautifulSoup?

Real-world use cases are often more complex than simply retrieving a title or a link. You'll need to handle the’structured data extraction Like tables and lists, the page numbers, and common scraping errors.
Extract tables and lists
Websites often present their data in HTML tables (<table>, <tr>, <th>, <td>) or lists (/ with ). To convert these structures into usable data, you need to iterate through them line by line or element by element.
To extract a HTML tablethe principle is simple:
- Retrieve the headers (
<th>) to identify column headings. - Go through each line (
<tr>) and search for cells (<td>) that contain the data. - Store the information in a list or a dictionary.
For a HTML list ( Where ) :
- Find all the tags
with find_all. - Retrieve their content (text or link) and add it to a Python list.
Here's an example with a table:
html = """
<table>
<tr>
<th>Last name</th>
<th>Age</th>
<th>Town</th>
</tr>
<tr>
<td>Alice</td>
<td>25</td>
<td>Paris</td>
</tr>
<tr>
<td>Bob</td>
<td>30</td>
<td>Lyon</td>
</tr>
</table>
"""
# Create BeautifulSoup object
soup = BeautifulSoup(html, "html.parser")
# Extract headers from array
headers = [th.get_text(strip=True) for th in soup.find_all("th")]
print("Headers:", headers)
# Extract data rows (skip 1st row as these are the headers)
rows = []
for tr in soup.find_all("tr")[1:]:
cells = [td.get_text(strip=True) for td in tr.find_all("td")]
if cells:
rows.append(cells)
print("Lines :", rows)
Here, find_all("th") retrieves the headings and find_all("td") retrieves the cells from each row. You loop through the <tr> to rebuild the table row by row.
Here's an example on a list:
from bs4 import BeautifulSoup
html_list = """
- Apple
- Banana
- Orange
"""
soup = BeautifulSoup(html_list, "html.parser")
# Retrieve list items
items = [li.get_text(strip=True) for li in soup.find_all("li")]
print("Extracted list:", items) # ["Apple", "Banana", "Orange"]
Each is directly converted into Python list element, which gives the result ["Apple", "Banana", "Orange"].
Manage pagination and links
Often, the data doesn't fit on a single page. It is distributed via links to “next page” or a numbered pagination (?page=1, ?page=2, etc.). In both cases, you must curl to retrieve all the pages and merge the data.
Example using a page parameter:
import time
import requests
from bs4 import BeautifulSoup
# Example of URL with pagination
BASE_URL = "https://bonjour.com/articles?page={}"
HEADERS = {"User-Agent": "Mozilla/5.0"}
all_articles = []
# Assume 5 pages to browse
for page in range(1, 6):
url = BASE_URL.format(page)
r = requests.get(url, headers=HEADERS, timeout=20)
if r.status_code == 200:
soup = BeautifulSoup(r.text, "html.parser")
# Extract article titles
articles = [h2.get_text(strip=True) for h2 in soup.find_all("h2", class_="title")]
all_articles.extend(articles)
else:
print(f "Error on page {page} (code : {r.status_code})")
time.sleep(1.0) # politeness
print("Articles retrieved :", all_articles)
A brief step-by-step explanation:
- Prepare the’URLs with a spot
{}to insert the page number.
BASE_URL = "https://bonjour.com/articles?page={}
- Some websites block requests that do not include a browser ID. Add a User-Agent Avoid being mistaken for a bot.
headers = {"User-Agent": "Mozilla/5.0"}
requests.get(url, headers=headers)
- Page Footer 1 to 5.
for page in range(1, 6):
- Retrieve the HTML from the page using requests.
requests.get(url)
- Limits the wait time if the site does not respond (20-second timeout).
requests.get(url, timeout=20)
- Analyze the page with BeautifulSoup.
BeautifulSoup(response.text, "html.parser")
- Retrieves all the article titles.
find_all("h2", class_="title")
- Adds the items found to a master list.
all_articles.extend(articles)
- Enter a break between each request to avoid overloading the server and getting banned.
time.sleep(1.0)
After the loop, all_articles contains all 5 page titles.
Common mistakes and challenges
Web scraping isn't just about pressing a button. You'll run into some common obstacles:
- HTTP errors : 404 (page not found), 403 (no entry) and 500 (server-side error).
Management example:
response = requests.get(url)
if response.status_code == 200:
print("Page retrieved successfully")
elif response.status_code == 404:
print("Error: Page not found")
else:
print("Status code returned:", response.status_code)
- Sites that block scraping : Some detect automated requests and block access.
- Dynamic pages (JavaScript) : BeautifulSoup only reads the Static HTML. If the page loads its content using JavaScript, you won't see anything. In that case, you should use a tool like Selenium Where Playwright.
To scrape effectively without getting blocked or damaging the site, here are the best practices:
- Respect the robots.txt file of the website you want to analyze.
- Set up some deadlines between queries with
time.sleep()so as not to overload the server. - Uses proxies and spin them.
- Change your User-Agent.
Web scraping with Selenium and BeautifulSoup?

Selenium control a real browser, runs the JavaScript and displays the page as if a human were browsing it. BeautifulSoup Then analyze the HTML once the page has fully loaded, and extract all the data you want.
Step 1: Install Selenium and BeautifulSoup
Here, instead of requests, you use Selenium to retrieve the page's content. Install the two libraries using this command:
pip install selenium beautifulsoup4
Next, you need to download a WebDriver tailored to your browser version. For Google Chrome, it's ChromeDriver. You have two options: place it in the same folder as your Python script, or’add to the PATH environment variable of your system.
Step 2: Configure Selenium
Start by importing webdriver using Selenium to control the browser.
from selenium import webdriver
from selenium.webdriver.common.by import By
Next, launch a browser. The browser opens the page and runs the JavaScript, in this case Chrome.
driver = webdriver.Chrome()
Tells the browser which URLs visit.
driver.get("https://www.exemple.com")
If the page takes a while to display certain elements, you can tell Selenium to wait up to 10 seconds.
driver.implicitly_wait(10)
Step 3: Retrieving page content
Once the page has loaded, you retrieve the Full DOM : the HTML source code generated after the JavaScript is executed.
html_content = driver.page_source
Step 4: HTML analysis with BeautifulSoup
Now pass this source code to BeautifulSoup To use it:
from bs4 import BeautifulSoup
# Create a BeautifulSoup object
soup = BeautifulSoup(html_content, 'html.parser')
# Example: Retrieve all headings from the page
headings = soup.find_all('h2')
for heading in headings:
print(heading.get_text())
BeautifulSoup offers powerful methods such as find(), find_all() and CSS selectors for target and extract specific HTML elements.
Step 5: Closing the browser
Very important: Always close your browser after running the program to free up resources.
driver.quit()
This way, you combine the power of Selenium to simulate human navigation (clicks, scrolls) combined with BeautifulSoup's efficiency in parsing HTML. This is the ideal method for scraping a site that loads its data via JavaScript.
FAQs
What's the best tool for web scraping in Python?
There is no’scraping tool There's no one-size-fits-all solution—rather, solutions tailored to your project. Here are the three most commonly used ones:
- BeautifulSoup : Simple and effective for parsing HTML and quickly extracting content. Ideal if you're just starting out or working on small projects.
- Scrapy : a comprehensive framework designed to manage large data volumes with advanced features.
- Playwright : Perfect for complex JavaScript-driven websites. It simulates a real browser and interacts with the page just like a human would.
How do I use BeautifulSoup to extract the content of a div tag?
With BeautifulSoup, you target a specific tag using a CSS selector. To extract the content of a tag <div>, follow these steps.
1. Retrieve the web page using requests, then parse the HTML with BeautifulSoup.
from bs4 import BeautifulSoup
import requests
url = "URL_DE_TON_SITE" # Remplace par l'URL réelle
reponse = requests.get(url)
html_content = reponse.text
soup = BeautifulSoup(html_content, "html.parser")
2. Use the method select() using your CSS selector to target the tag <div>.
To retrieve the first element, use soup.select_one. To retrieve all the elements, use soup.select. Here is an example of HTML code:
<div>
<h2>Titre de l'article</h2>
<p>Voici le contenu du paragraphe.</p>
</div>
And here's an example using the CSS selector div.article :
# Récupérer le premier div avec la classe "article"
div_article = soup.select_one("div.article")
# Afficher son contenu texte
if div_article:
print(div_article.get_text(strip=True))
3. Remove the items inside the <div>.
# Récupérer le titre à l'intérieur du div
titre = soup.select_one("div.article h2").get_text()
# Récupérer le paragraphe à l'intérieur du div
paragraphe = soup.select_one("div.article p").get_text()
print("Titre :", titre)
print("Paragraphe :", paragraphe)
How do I use Requests and BeautifulSoup together?
These two libraries are additional : One downloads the page, the other analyzes it.
1. requests sends an HTTP request and downloads the raw HTML code of the page.
import requests
url = "https://sitecible.com"
response = requests.get(url) # requête HTTP
print(response.text) # affiche le HTML brut
At this point, you just have a huge block of text full of tags (<html>, <div>, <p>etc.).
2. BeautifulSoup analyzes this raw HTML and converts it into a organized structure. You can then browse the page, locate the tags, and retrieve the data.
from bs4 import BeautifulSoup
soup = BeautifulSoup(response.text, "html.parser") # analyse le HTML
titre = soup.find("h1").get_text() # extrait le contenu d'un <h1>
print(titre)
Why doesn't my web scraping code work on some sites?
Sometimes your script won't retrieve anything. Some sites don't provide all their content directly in HTML: they use JavaScript to load the data dynamically. However, BeautifulSoup cannot parse the data returned by JavaScript. In that case, you should turn to tools like Playwright Where Selenium.
What role does BeautifulSoup play in web scraping?
BeautifulSoup plays the role of’HTML parser. It converts a page's code into a structured object that you can easily navigate. In short, it's the a converter between raw HTML and your Python code.
Web Scraping: BeautifulSoup or Scrapy?
BeautifulSoup and Scrapy are really different, even though both are used for web scraping.
| BeautifulSoup | Scrapy |
|---|---|
| A simple library for parsing HTML and extracting data. | A comprehensive framework that handles the entire scraping process: requests, link tracking, pagination, data export, and error handling. |
BeautifulSoup is the perfect choice for beginners: it makes the’HTML data extraction in Python Quick and easy.
If you'd rather keep the coding to a minimum, the tool Bright Data offers ready-to-use datasets to reduce the amount of coding required.





