Skip to content

Repository files navigation

📰 ElPaisScraper

ElPaisScraper is a Java-based automation project that uses Selenium WebDriver, MyMemory Translation API, and BrowserStack for web scraping, translation, and cross-browser testing.


Features

  • Scrapes the first 5 articles from the Opinion section of El País
  • Extracts title, content, and downloads images
  • Translates the article titles from Spanish to English using the MyMemory API
  • Analyzes repeated words in the translated titles
  • Performs parallel cross-browser testing using BrowserStack
  • Saves article images locally

Technologies Used

  • Java (JDK 20+)
  • Selenium WebDriver
  • MyMemory Translation API
  • Apache HttpClient & Gson
  • BrowserStack Automate
  • Maven (for dependency management)

Project Structure

ElPaisScraper/ ├── pom.xml ├── images ├── /src/ │ └── /main/ │ └── /java/ │ └── com/akshay/ │ ├── Article.java │ ├── BrowserStackTest.java │ ├── ElPaisScraper.java │ ├── LibreTranslateutil.java │ └── Main.java └── [downloaded article images]


How to Run

1. Clone the Repository

git clone https://github.com/your-username/ElPaisScraper.git cd ElPaisScraper

  1. Import the project into IntelliJ IDEA or Eclipse Make sure to use a JDK version 17 or above and enable Maven support.

  2. Add BrowserStack credentials Update the code with your actual browserstack.username and browserstack.access_key.

Alternatively, you can use environment variables if preferred.

  1. Run the project First, run ElPaisScraper.java to scrape articles and download images

Then, run BrowserStackTest.java to perform parallel browser testing

Output

Article titles and content printed to console

Article images downloaded to the images/ folder

Repeated words from translated titles printed

Parallel browser session logs viewable on BrowserStack Dashboard


3. pom.xml Verification
Make sure it contains all necessary dependencies:

org.seleniumhq.selenium selenium-java 4.21.0 org.apache.httpcomponents.client5 httpclient5 5.3.1 com.google.code.gson gson 2.10.1

Notes Ensure your internet connection is active when running the scraper and translation API

For BrowserStack testing, free accounts are limited — plan usage accordingly

This project demonstrates a mix of automation, web scraping, API integration, and testing — all in one place

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages