Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

Β 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

πŸš€ High-Performance Multi-threaded Web Scraper

A professional-grade web scraping solution with dynamic proxy rotation and asynchronous data processing, designed for high-concurrency financial data collection.

✨ Features

  • πŸ”„ Dynamic Proxy Rotator: Intelligent proxy selection based on health scores with automatic recovery
  • ⚑ Asyncio + ThreadPoolExecutor: Optimal combination for I/O-bound and CPU-bound tasks
  • πŸ”’ Thread-Safe Buffer: RLock-protected storage with MD5-based deduplication
  • βœ… JSON Schema Validation: Comprehensive validation for financial data structures
  • πŸ“Š Zero Data Loss: Guaranteed data integrity under 100+ concurrent requests
  • 🎯 Smart Retry Logic: Automatic failover with alternative proxies

πŸ—οΈ Architecture

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                    FinancialDataScraper                      β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚  β”‚ DynamicProxy     β”‚    β”‚ ThreadSafeBuffer             β”‚   β”‚
β”‚  β”‚ Rotator          β”‚    β”‚ - RLock Protection           β”‚   β”‚
β”‚  β”‚ - Health Scores  β”‚    β”‚ - MD5 Deduplication          β”‚   β”‚
β”‚  β”‚ - Weighted Pick  β”‚    β”‚ - Statistics Tracking        β”‚   β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β”‚                                                              β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚  β”‚              asyncio Event Loop                       β”‚   β”‚
β”‚  β”‚  (I/O-bound: HTTP Requests via aiohttp)              β”‚   β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β”‚                          ⬇️                                  β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚  β”‚           ThreadPoolExecutor                         β”‚   β”‚
β”‚  β”‚  (CPU-bound: JSON Schema Validation)                 β”‚   β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

πŸ“¦ Installation

# Clone the repository
git clone https://github.com/yourusername/multi-threaded-web-scraper.git
cd multi-threaded-web-scraper

# Install dependencies
pip install aiohttp

πŸš€ Quick Start

Basic Usage

import asyncio
from web_scraper import FinancialDataScraper, FINANCIAL_ENDPOINTS, PROXY_POOL

async def main():
    scraper = FinancialDataScraper(
        endpoints=FINANCIAL_ENDPOINTS,
        proxies=PROXY_POOL,
        max_workers=10
    )
    
    # Run with 100 concurrent requests
    stats = await scraper.run(request_count=100)
    
    print(f"Stored: {stats['stored']}")
    print(f"Duplicates: {stats['duplicates']}")

if __name__ == "__main__":
    asyncio.run(main())

Custom Configuration

# Define your own endpoints
endpoints = [
    "https://api.example.com/stock/AAPL",
    "https://api.example.com/stock/GOOGL",
    "https://api.example.com/stock/MSFT"
]

# Define your proxy pool
proxies = [
    "http://proxy1.example.com:8080",
    "http://proxy2.example.com:8080",
    "http://proxy3.example.com:8080"
]

scraper = FinancialDataScraper(
    endpoints=endpoints,
    proxies=proxies,
    max_workers=20  # Adjust thread pool size
)

πŸ§ͺ Running Benchmarks

Test the scraper with 100 concurrent requests:

python benchmark_test.py

Expected Output

======================================================================
COMPREHENSIVE BENCHMARK TEST
======================================================================

Test Configuration:
  - Concurrent Requests: 100
  - Financial Endpoints: 5
  - Expected Records: 500

============================================================
BENCHMARK RESULTS
============================================================
Number of Requests:        100
Total Expected Records:    500
Total Stored:              500
Data Loss:                 0
Success Rate:              100.00%
Elapsed Time:              0.280 seconds
Requests/Second:           1788.84
============================================================
βœ… SUCCESS: Zero data loss confirmed!

πŸ“ Project Structure

multi-threaded-web-scraper/
β”œβ”€β”€ web_scraper.py          # Main scraper implementation
β”œβ”€β”€ benchmark_test.py       # Comprehensive benchmark tests
β”œβ”€β”€ README.md               # This file
β”œβ”€β”€ LICENSE                 # MIT License
└── requirements.txt        # Python dependencies

πŸ”§ Components

1. DynamicProxyRotator

Intelligent proxy management with health-based scoring:

  • Health Score: Tracks success/failure ratio (0.0 - 1.0)
  • Weighted Selection: Probabilistic selection favoring healthy proxies
  • Automatic Recovery: Failed proxies can recover over time
@dataclass
class ProxyHealth:
    url: str
    health_score: float = 1.0
    failures: int = 0
    
    def record_success(self):
        self.health_score = min(1.0, self.health_score + 0.1)
        
    def record_failure(self):
        self.health_score = max(0.0, self.health_score - 0.2)

2. ThreadSafeBuffer

Thread-safe data storage with deduplication:

  • RLock Protection: Prevents race conditions
  • MD5 Hashing: Efficient duplicate detection
  • Statistics Tracking: Monitors received, stored, and lost data

3. FinancialDataScraper

Main orchestrator combining asyncio and threading:

  • Async HTTP Requests: Non-blocking I/O with aiohttp
  • CPU-bound Validation: Offloaded to ThreadPoolExecutor
  • Retry Logic: Automatic failover with alternative proxies

🎯 Performance Characteristics

Metric Value
Concurrent Requests 100+
Data Loss 0%
Throughput ~1800 req/sec
Execution Time < 1 second
Memory Efficiency O(n) with deduplication

πŸ” Thread Safety

The implementation ensures thread safety through:

  1. RLock for buffer operations
  2. Atomic operations for proxy health updates
  3. Proper executor shutdown on completion
  4. Session management with async context managers

πŸ“ License

MIT License - see LICENSE for details.

🀝 Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

  1. Fork the repository
  2. Create your feature branch (git checkout -b feature/amazing-feature)
  3. Commit your changes (git commit -m 'Add amazing feature')
  4. Push to the branch (git push origin feature/amazing-feature)
  5. Open a Pull Request

πŸ“§ Support

For issues and questions, please open an issue on GitHub.


Built with ❀️ using Python, asyncio, and ThreadPoolExecutor

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages