Skip to content
RyankhlifiPublic

About

A Spring Boot web crawler that stores URLs persistently and powers fast searches using an inverted index in MongoDB. Designed for multi-threaded crawling and efficient link discovery

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

37 Commits

Folders and files

Repository files navigation

JC - Java Web Crawler

A multi-threaded web crawler built with Spring Boot that crawls web pages, extracts content, and stores indexed data in MongoDB.

Features

  • Multi-threaded Crawling: Uses a configurable thread pool to crawl multiple pages concurrently
  • URL Discovery: Automatically discovers and normalizes links from crawled pages
  • Content Indexing: Extracts and tokenizes page content using Apache Lucene's Standard Analyzer
  • MongoDB Storage: Persists crawled web pages with their URLs and processed content
  • Duplicate Detection: Tracks visited URLs to avoid re-crawling the same pages

Tech Stack

  • Java 24
  • Spring Boot 4.0.1
  • MongoDB - Document storage for web archives
  • Jsoup - HTML parsing and web page fetching
  • Apache Lucene - Text analysis and tokenization
  • Lombok - Reduces boilerplate code

Project Structure

src/main/java/org/rayen/jc/
├── JcApplication.java              # Spring Boot application entry point
├── Configuration/
│   └── Config.java                 # MongoDB converter configuration
├── CrawlerEngine/
│   ├── Crawler/
│   │   ├── Crawler.java            # Main crawler logic with multi-threading
│   │   └── Discovery.java          # URL frontier and visited set management
│   └── Indexer/
│       └── IndexerService.java     # Content normalization and storage
├── Entity/
│   └── WebArchive.java             # MongoDB document model
├── Repository/
│   └── WebArchiveRepository.java   # MongoDB repository interface
└── Runner/
    └── CrawlRunner.java            # CommandLineRunner to start crawling

Prerequisites

  • Java 24
  • Maven 3.x
  • MongoDB instance

Configuration

Set the following environment variables:

export MONGO_URI=mongodb://localhost:27017
export MONGO_DB=jc_database
export CRAWLER_THREADPOOL_SIZE=4  # Optional: number of crawler threads (default: 2)

Or configure in src/main/resources/application.properties:

spring.application.name=JC
spring.mongodb.uri=${MONGO_URI}
spring.mongodb.database=${MONGO_DB}

# Crawler configuration
crawler.threadpool.size=${CRAWLER_THREADPOOL_SIZE:2}

Configuring the Crawler Thread Pool

The crawler uses a configurable thread pool to control the number of concurrent crawling threads. You can adjust this based on your system resources and target server capacity.

Option 1: Environment Variable

export CRAWLER_THREADPOOL_SIZE=4

Option 2: Application Properties

crawler.threadpool.size=4

Option 3: Command Line

./mvnw spring-boot:run -Dcrawler.threadpool.size=4

Recommended Values:

  • 2 threads (default): Conservative, suitable for most use cases
  • 4-8 threads: Higher throughput for faster systems with good network
  • 1 thread: Minimum impact on target servers, useful for polite crawling

Building

./mvnw clean package

Running

./mvnw spring-boot:run

The crawler will start with a default seed URL (https://www.gatech.edu) and crawl up to 5 pages.

How It Works

  1. Initialization: The CrawlRunner starts the crawler with a seed URL
  2. URL Frontier: The Discovery class manages a blocking queue of URLs to visit
  3. Crawling: Worker threads fetch pages using Jsoup, extract links, and add them to the frontier
  4. Indexing: The IndexerService tokenizes page content (up to 2000 tokens) using Lucene's Standard Analyzer
  5. Storage: Processed content and URLs are saved to MongoDB as WebArchive documents

Testing

./mvnw test

Trade-offs

Design Decisions and Their Implications

Decision Benefit Trade-off
In-memory URL frontier Fast URL access, simple implementation Not persistent across restarts; limited by available memory
Configurable thread pool (default 2 threads) Controlled resource usage, configurable via properties Requires restart to change; user must balance throughput vs. server load
ConcurrentHashMap for visited URLs Thread-safe duplicate detection Memory grows with crawled URLs; no persistence
Blocking queue for URL frontier Natural backpressure, thread coordination Threads block when queue is empty
HTTPS-only filtering Ensures secure connections Excludes HTTP-only content
Token limit (2000 tokens) Bounded storage per document, consistent processing time May truncate content from large pages
Synchronous MongoDB writes Simple, consistent data storage Slower than batched/async writes
Single seed URL Simple initialization Requires restart to crawl different domains
4-second connection timeout Prevents hanging on slow servers May miss content from high-latency servers

About

A Spring Boot web crawler that stores URLs persistently and powers fast searches using an inverted index in MongoDB. Designed for multi-threaded crawling and efficient link discovery

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages