ASSISTANT
Certainly! Below is a Python script that accomplishes the following tasks:
1. **Parses the List of URLs:** It processes the provided list of Fiverr URLs.
2. **Fetches Each Webpage:** Utilizes the `requests` library to send HTTP requests to each URL.
3. **Parses the HTML Content:** Uses `BeautifulSoup` from the `bs4` library to parse the HTML content of each page.
4. **Extracts Specific Sections:** Searches for specific HTML tags or sections related to your needs (e.g., service descriptions, tags, etc.).
5. **Saves Extracted Data to a File:** Writes the extracted sections into a text file for further analysis or review.
### **Prerequisites:**
Before running the script, ensure you have the necessary Python libraries installed. You can install them using `pip`:
```bash
pip install requests beautifulsoup4
```
### **Script Explanation:**
- **URL List:** The script expects the list of URLs to be provided in a text file named `urls.txt`, with each URL on a separate line. However, if you prefer, you can define the URLs directly within the script.
- **Target Tags:** By default, the script is set to extract the service description section based on Fiverr's typical HTML structure. You may need to adjust the HTML tags or identifiers (`class`, `id`, etc.) based on the actual structure of the Fiverr pages you're targeting.
- **Error Handling:** The script includes basic error handling to manage failed HTTP requests or parsing issues gracefully.
### **Python Script:**
```python
import requests
from bs4 import BeautifulSoup
import time
# Function to read URLs from a text file
def read_urls(file_path):
with open(file_path, 'r') as file:
urls = file.read().splitlines()
# Filter out any empty lines
urls = [url.strip() for url in urls if url.strip()]
return urls
# Function to extract desired section from the webpage
def extract_section(html_content):
soup = BeautifulSoup(html_content, 'html.parser')
# Example: Extract the service description
# Note: The actual HTML structure of Fiverr pages may vary.
# You may need to inspect a Fiverr page to adjust the selectors accordingly.
description_section = soup.find('div', class_='description') # Example class name
if description_section:
return description_section.get_text(strip=True)
else:
return "Description section not found."
# Function to process each URL
def process_urls(urls, output_file):
headers = {
'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) ' \
'AppleWebKit/537.36 (KHTML, like Gecko) ' \
'Chrome/90.0.4430.93 Safari/537.36'
}
with open(output_file, 'w', encoding='utf-8') as f_out:
for idx, url in enumerate(urls, 1):
print(f"Processing ({idx}/{len(urls)}): {url}")
try:
response = requests.get(url, headers=headers, timeout=10)
response.raise_for_status() # Raise HTTPError for bad responses
section_text = extract_section(response.text)
# Write to file
f_out.write(f"URL {idx}: {url}\n")
f_out.write(f"Extracted Section:\n{section_text}\n")
f_out.write("="*80 + "\n\n")
# Optional: Wait between requests to be polite
time.sleep(1)
except requests.exceptions.RequestException as e:
print(f"Failed to fetch {url}: {e}")
f_out.write(f"URL {idx}: {url}\n")
f_out.write(f"Error: {e}\n")
f_out.write("="*80 + "\n\n")
except Exception as e:
print(f"An error occurred while processing {url}: {e}")
f_out.write(f"URL {idx}: {url}\n")
f_out.write(f"Unexpected Error: {e}\n")
f_out.write("="*80 + "\n\n")
# Main function
def main():
# Option 1: Read URLs from a file named 'urls.txt'
# Each URL should be on a separate line
url_file = 'urls.txt'
try:
urls = read_urls(url_file)
if not urls:
print(f"No URLs found in {url_file}. Please add URLs, one per line.")
return
except FileNotFoundError:
print(f"File {url_file} not found. Please create the file with URLs, one per line.")
return
# Option 2: Define URLs directly in the script
# Uncomment the following lines and comment out the Option 1 block if preferred
"""
urls = [
"https://www.fiverr.com/ryancanadadev/create-openai-or-gpt3-or-chatgpt-or-machine-learning-app-for-you?context_referrer=subcategory_listing&...",
"https://www.fiverr.com/nachtigal4/build-a-prediction-model-f96a?context_referrer=subcategory_listing&...",
# Add all other URLs here
]
"""
output_file = 'extracted_sections.txt'
process_urls(urls, output_file)
print(f"Extraction complete. Check the '{output_file}' file for results.")
if __name__ == "__main__":
main()
```
### **Instructions to Run the Script:**
1. **Prepare the URL List:**
- Create a text file named `urls.txt` in the same directory as your Python script.
- Paste each Fiverr URL on a separate line. For example:
```
https://www.fiverr.com/ryancanadadev/create-openai-or-gpt3-or-chatgpt-or-machine-learning-app-for-you?context_referrer=...
https://www.fiverr.com/nachtigal4/build-a-prediction-model-f96a?context_referrer=...
...
```
2. **Adjust the Extraction Logic (If Necessary):**
- The current script attempts to find a `<div>` with the class `description`. Depending on Fiverr's actual HTML structure, you might need to inspect the page and adjust the selector in the `extract_section` function. For example:
```python
# Example adjustment
description_section = soup.find('div', class_='gig-description')
```
- Use your browser's developer tools (usually accessed by pressing `F12`) to inspect the elements you wish to extract and modify the selector accordingly.
3. **Run the Script:**
Execute the script using Python:
```bash
python your_script_name.py
```
Replace `your_script_name.py` with the actual name of your Python script file.
4. **View the Results:**
- After the script completes, open the `extracted_sections.txt` file to review the extracted content.
- Each entry will include the URL, the extracted section (or an error message if extraction failed), and a separator line for clarity.
### **Notes and Best Practices:**
- **Respect Website's Terms of Service:** Ensure that web scraping Fiverr's website complies with their [Terms of Service](https://www.fiverr.com/terms_of_service). Unauthorized scraping may lead to your IP being blocked or other legal consequences.
- **Handle Rate Limiting:** The script includes a `time.sleep(1)` to pause for 1 second between requests, which helps in preventing overwhelming the server and reduces the risk of being rate-limited or blocked. You can adjust or remove this delay based on your needs and Fiverr's policies.
- **Dynamic Content:** Some sections of Fiverr pages might be loaded dynamically using JavaScript, which `requests` and `BeautifulSoup` cannot handle. In such cases, consider using a headless browser like Selenium.
- **Error Handling:** The script includes basic error handling. For more robust applications, consider logging errors to a separate file and implementing retries for transient issues.
- **Extending Functionality:** Depending on your requirements, you can extend the script to extract additional information such as seller ratings, pricing, delivery time, etc., by adjusting the `extract_section` function accordingly.
### **Conclusion**
This script provides a foundational approach to scraping specific sections from a list of Fiverr URLs. Customize the extraction logic based on the exact HTML structure of the Fiverr pages you are targeting. Always ensure compliance with website policies and handle data responsibly.