Web Scraping | Remote | Python Experience

Posted 9/14/2026
Location United States, United States
Type Contract
Compensation US Dollars 25 / Hourly

Job description

**Location:** Remote (United States) **Duration:** September–December 2026 **Hours:** 15–20 hours per week **Overview** This internship offers the opportunity to build real datasets used by federal innovation and prize competitions. You will write and run Python extraction scripts against public sources including academic directories, lab pages, professional societies, conference programs, and open research APIs. The role involves delivering clean, verified CSV files with source URLs on every record, working on real datasets with real deadlines. **Key Responsibilities** - Write and run extraction scripts using BeautifulSoup, Scrapy, Selenium, or Playwright - Use public APIs and JSON endpoints where available, in preference to page scraping - Handle pagination, inconsistent markup, and varied formats without losing records - Deliver de-duplicated CSV or Excel output with a source URL on every row - Validate before delivery — no malformed rows, no duplicates, no guessed email addresses - Stay within each site's terms of use, robots.txt, and rate limits **Required Qualifications** - Current student or recent graduate in CS, data science, information science, or similar - Working Python, and familiarity with at least one extraction library - Comfortable reading API docs, handling JSON, and communicating about blockers - Careful and precise — accuracy matters more here than speed or clever code - Must be authorized to work in the United States without employer sponsorship, now or in the future **Work Location & Requirements** - Fully remote; must reside in the United States - Part-time internship position

Work setting

Remote