A multi-threaded web crawler built with Spring Boot that crawls web pages, extracts content, and stores indexed data in MongoDB.
- Multi-threaded Crawling: Uses a configurable thread pool to crawl multiple pages concurrently
- URL Discovery: Automatically discovers and normalizes links from crawled pages
- Content Indexing: Extracts and tokenizes page content using Apache Lucene's Standard Analyzer
- MongoDB Storage: Persists crawled web pages with their URLs and processed content
- Duplicate Detection: Tracks visited URLs to avoid re-crawling the same pages
- Java 24
- Spring Boot 4.0.1
- MongoDB - Document storage for web archives
- Jsoup - HTML parsing and web page fetching
- Apache Lucene - Text analysis and tokenization
- Lombok - Reduces boilerplate code
src/main/java/org/rayen/jc/
├── JcApplication.java # Spring Boot application entry point
├── Configuration/
│ └── Config.java # MongoDB converter configuration
├── CrawlerEngine/
│ ├── Crawler/
│ │ ├── Crawler.java # Main crawler logic with multi-threading
│ │ └── Discovery.java # URL frontier and visited set management
│ └── Indexer/
│ └── IndexerService.java # Content normalization and storage
├── Entity/
│ └── WebArchive.java # MongoDB document model
├── Repository/
│ └── WebArchiveRepository.java # MongoDB repository interface
└── Runner/
└── CrawlRunner.java # CommandLineRunner to start crawling
- Java 24
- Maven 3.x
- MongoDB instance
Set the following environment variables:
export MONGO_URI=mongodb://localhost:27017
export MONGO_DB=jc_database
export CRAWLER_THREADPOOL_SIZE=4 # Optional: number of crawler threads (default: 2)Or configure in src/main/resources/application.properties:
spring.application.name=JC
spring.mongodb.uri=${MONGO_URI}
spring.mongodb.database=${MONGO_DB}
# Crawler configuration
crawler.threadpool.size=${CRAWLER_THREADPOOL_SIZE:2}The crawler uses a configurable thread pool to control the number of concurrent crawling threads. You can adjust this based on your system resources and target server capacity.
Option 1: Environment Variable
export CRAWLER_THREADPOOL_SIZE=4Option 2: Application Properties
crawler.threadpool.size=4Option 3: Command Line
./mvnw spring-boot:run -Dcrawler.threadpool.size=4Recommended Values:
- 2 threads (default): Conservative, suitable for most use cases
- 4-8 threads: Higher throughput for faster systems with good network
- 1 thread: Minimum impact on target servers, useful for polite crawling
./mvnw clean package./mvnw spring-boot:runThe crawler will start with a default seed URL (https://www.gatech.edu) and crawl up to 5 pages.
- Initialization: The
CrawlRunnerstarts the crawler with a seed URL - URL Frontier: The
Discoveryclass manages a blocking queue of URLs to visit - Crawling: Worker threads fetch pages using Jsoup, extract links, and add them to the frontier
- Indexing: The
IndexerServicetokenizes page content (up to 2000 tokens) using Lucene's Standard Analyzer - Storage: Processed content and URLs are saved to MongoDB as
WebArchivedocuments
./mvnw test| Decision | Benefit | Trade-off |
|---|---|---|
| In-memory URL frontier | Fast URL access, simple implementation | Not persistent across restarts; limited by available memory |
| Configurable thread pool (default 2 threads) | Controlled resource usage, configurable via properties | Requires restart to change; user must balance throughput vs. server load |
| ConcurrentHashMap for visited URLs | Thread-safe duplicate detection | Memory grows with crawled URLs; no persistence |
| Blocking queue for URL frontier | Natural backpressure, thread coordination | Threads block when queue is empty |
| HTTPS-only filtering | Ensures secure connections | Excludes HTTP-only content |
| Token limit (2000 tokens) | Bounded storage per document, consistent processing time | May truncate content from large pages |
| Synchronous MongoDB writes | Simple, consistent data storage | Slower than batched/async writes |
| Single seed URL | Simple initialization | Requires restart to crawl different domains |
| 4-second connection timeout | Prevents hanging on slow servers | May miss content from high-latency servers |