Bulk URL Checker with uv: Find Broken Links in Python
Learn how to build a bulk URL checker in Python with uv: check hundreds of URLs concurrently, flag broken links, 404s, and timeouts, and schedule it in CI.
Updated Published 43 min read

Broken links hurt your SEO, annoy visitors, and make a site look abandoned. I wrote this bulk URL checker to scan hundreds of URLs at once, flag 404s, 500s, timeouts, and connection errors, then write the broken ones to a file for review. It runs as a single Python script with uv run. No virtualenv, no requirements.txt, no pip install. Updated for uv 0.12 (Aug 2026).
New to uv?
If you’re new to uv or want to learn how to set up full Python projects, start with Getting Started with uv: Python Project Setup in 2026 before diving into this script.
What this script does
- Checks multiple URLs concurrently with ThreadPoolExecutor
- Auto-prepends
https://when no scheme is provided - Classifies errors: TIMEOUT, CONNECTION_ERROR, SSL_ERROR, CLIENT_ERROR, SERVER_ERROR
- Reads URLs from a text file (skips comments and blank lines)
- Writes broken URLs to an output file for review
- Real-time progress counter while checking
- CLI flags for automation:
--timeout,--workers,--output,--head,--insecure - Exits with code 1 when broken links are found (for CI integration)
Why uv? Run Python scripts with zero setup
The script uses PEP 723 inline script metadata, a # /// script block at the top of the file that declares dependencies. When you run uv run url_checker.py, uv reads that block, installs requests in an isolated environment, and executes the script. No virtualenv activation, no requirements.txt, no pip install.
You can pin the Python version with requires-python, pin dependency versions (requests<3), and lock the whole resolution with uv lock --script url_checker.py to get a .lock file for reproducible runs. The same pattern powers other single-file tools. See uv run: run Python scripts with zero dependency management for the full workflow, and the same pattern powers our uv text-to-speech script as another example.
The script
Save this as url_checker.py:
#!/usr/bin/env -S uv run --script
# /// script
# requires-python = ">=3.10"
# dependencies = [
# "requests<3",
# ]
# [tool.uv]
# exclude-newer = "2026-08-01T00:00:00Z"
# ///
"""
Bulk URL checker — checks hundreds of URLs concurrently,
flags broken links (4xx, 5xx, timeouts, SSL errors), and
writes results to a file.
Usage:
uv run url_checker.py urls.txt
uv run url_checker.py urls.txt --workers 20 --timeout 15
uv run url_checker.py urls.txt --head --output broken.txt
"""
import argparse
import csv
import sys
import time
from concurrent.futures import ThreadPoolExecutor, as_completed
from urllib.parse import urlparse
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
# Shared session with connection pooling and a realistic User-Agent
SESSION = requests.Session()
SESSION.headers["User-Agent"] = "url-checker/2.0 (bitdoze.com)"
retries = Retry(total=1, backoff_factor=0.5, status_forcelist=[429, 500, 502, 503, 504])
SESSION.mount("https://", HTTPAdapter(max_retries=retries))
SESSION.mount("http://", HTTPAdapter(max_retries=retries))
def normalize_url(url: str) -> str:
"""Add https:// if no scheme is present."""
if not url.startswith(("http://", "https://")):
return "https://" + url
return url
def check_url(url: str, timeout: float = 10, use_head: bool = False, verify: bool = True) -> dict:
"""
Check a single URL and return a result dict.
Returns dict with keys: url, status, status_code, error_type, response_time
"""
url = normalize_url(url)
start = time.monotonic()
methods = ("head", "get") if use_head else ("get",)
for method in methods:
try:
r = SESSION.request(
method, url, timeout=timeout, allow_redirects=True, verify=verify
)
elapsed = round(time.monotonic() - start, 2)
# HEAD rejected — fall back to GET
if r.status_code == 405 and method == "head":
continue
code = r.status_code
if 200 <= code < 300:
return {"url": url, "status": "OK", "status_code": code, "error_type": None, "response_time": elapsed}
elif 400 <= code < 500:
return {"url": url, "status": "CLIENT_ERROR", "status_code": code, "error_type": f"{code} Client Error", "response_time": elapsed}
else:
return {"url": url, "status": "SERVER_ERROR", "status_code": code, "error_type": f"{code} Server Error", "response_time": elapsed}
except requests.exceptions.SSLError as e:
return {"url": url, "status": "SSL_ERROR", "status_code": None, "error_type": f"SSL error: {str(e)[:80]}", "response_time": round(time.monotonic() - start, 2)}
except requests.exceptions.Timeout:
return {"url": url, "status": "TIMEOUT", "status_code": None, "error_type": "Connection timeout", "response_time": round(time.monotonic() - start, 2)}
except requests.exceptions.ConnectionError as e:
return {"url": url, "status": "CONNECTION_ERROR", "status_code": None, "error_type": f"Connection error: {str(e)[:80]}", "response_time": round(time.monotonic() - start, 2)}
except requests.exceptions.RequestException as e:
return {"url": url, "status": "ERROR", "status_code": None, "error_type": f"Request error: {str(e)[:80]}", "response_time": round(time.monotonic() - start, 2)}
return {"url": url, "status": "ERROR", "status_code": None, "error_type": "All methods failed", "response_time": round(time.monotonic() - start, 2)}
def read_urls(filename: str) -> list[str]:
"""Read URLs from a text file, one per line. Skips blank lines and # comments."""
urls = []
try:
with open(filename, encoding="utf-8") as f:
for line in f:
url = line.strip()
if url and not url.startswith("#"):
urls.append(url)
except FileNotFoundError:
print(f"Error: file '{filename}' not found.", file=sys.stderr)
sys.exit(1)
return urls
def check_urls(urls: list[str], timeout: float, workers: int, use_head: bool, verify: bool) -> list[dict]:
"""Check URLs concurrently and return a list of result dicts."""
results = []
with ThreadPoolExecutor(max_workers=workers) as pool:
futures = {pool.submit(check_url, u, timeout, use_head, verify): u for u in urls}
for i, future in enumerate(as_completed(futures), 1):
result = future.result()
results.append(result)
status_icon = "✓" if result["status"] == "OK" else "✗"
code_str = f" ({result['status_code']})" if result["status_code"] else ""
print(f" [{i}/{len(urls)}] {status_icon} {result['url']} — {result['status']}{code_str} ({result['response_time']}s)")
return results
def save_broken(results: list[dict], filename: str) -> None:
"""Write broken URLs grouped by error type."""
broken = [r for r in results if r["status"] != "OK"]
if not broken:
return
groups: dict[str, list[dict]] = {}
for r in broken:
groups.setdefault(r["status"], []).append(r)
with open(filename, "w", encoding="utf-8") as f:
f.write(f"# Broken URLs — checked {time.strftime('%Y-%m-%d %H:%M:%S')}\n\n")
for group_name in ("CLIENT_ERROR", "SERVER_ERROR", "TIMEOUT", "CONNECTION_ERROR", "SSL_ERROR", "ERROR"):
items = groups.get(group_name, [])
if items:
f.write(f"# {group_name}\n")
for r in items:
code = f" ({r['status_code']})" if r["status_code"] else ""
f.write(f"{r['url']}{code}\n")
f.write("\n")
def save_to_csv(results: list[dict], filename: str) -> None:
"""Export all results to CSV."""
with open(filename, "w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=["url", "status", "status_code", "error_type", "response_time"])
writer.writeheader()
writer.writerows(results)
def main() -> int:
parser = argparse.ArgumentParser(description="Bulk URL checker — find broken links fast")
parser.add_argument("file", nargs="?", default="urls.txt", help="file with URLs, one per line (default: urls.txt)")
parser.add_argument("--timeout", type=float, default=10, help="request timeout in seconds (default: 10)")
parser.add_argument("--workers", type=int, default=10, help="concurrent threads (default: 10)")
parser.add_argument("--output", default="problematic_urls.txt", help="output file for broken URLs")
parser.add_argument("--csv", default=None, help="export all results to a CSV file")
parser.add_argument("--head", action="store_true", help="use HEAD requests, fall back to GET on 405")
parser.add_argument("--insecure", action="store_true", help="disable TLS verification (internal hosts only)")
args = parser.parse_args()
if args.insecure:
print("WARNING: TLS verification disabled. Do not use --insecure on public URLs.", file=sys.stderr)
import urllib3
urllib3.disable_warnings(urllib3.exceptions.InsecureRequestWarning)
verify = not args.insecure
print(f"Reading URLs from '{args.file}'...")
urls = read_urls(args.file)
if not urls:
print("No URLs found.")
return 0
print(f"Found {len(urls)} URLs — checking with {args.workers} workers, {args.timeout}s timeout\n")
results = check_urls(urls, timeout=args.timeout, workers=args.workers, use_head=args.head, verify=verify)
ok = [r for r in results if r["status"] == "OK"]
broken = [r for r in results if r["status"] != "OK"]
print(f"\n{'=' * 50}")
print(f"RESULTS: {len(ok)} working, {len(broken)} broken out of {len(results)}")
print(f"{'=' * 50}")
if broken:
save_broken(results, args.output)
print(f"\nBroken URLs saved to '{args.output}'")
# Group summary
groups: dict[str, int] = {}
for r in broken:
groups[r["status"]] = groups.get(r["status"], 0) + 1
for name, count in sorted(groups.items()):
print(f" {name}: {count}")
if args.csv:
save_to_csv(results, args.csv)
print(f"All results exported to '{args.csv}'")
return 1 if broken else 0
if __name__ == "__main__":
sys.exit(main())How it works
Check the status code, not just exceptions
The original version of this script had a bug: it only caught exceptions. A 404 or 500 response doesn’t raise. requests.get() returns normally. So httpstat.us/404 was reported as “OK” in the sample output. For a tool whose whole purpose is finding broken links, that’s a problem.
The fix: classify on response.status_code. Only 2xx is OK. Everything else (4xx client errors, 5xx server errors) gets flagged as broken. If a URL redirects (301 to 404), allow_redirects=True means the script sees the final status code, so a redirect to a missing page is correctly flagged.
Common mistake
Checking only for exceptions misses 4xx and 5xx responses entirely. A 404 Not Found returns a normal response object. It does not raise. Always check response.status_code for the full picture.
Reuse connections with requests.Session()
Each bare requests.get() call opens a new TCP connection. The updated script creates a shared requests.Session() with connection pooling and HTTP keep-alive. On a list of 200+ URLs, this is noticeably faster because the script reuses existing connections to the same host instead of doing a fresh TCP handshake every time.
The session also sets a default User-Agent header and configures automatic retries for transient server errors (429, 500, 502, 503, 504) with exponential backoff.
HEAD requests first, GET fallback
With the --head flag, the script sends HEAD requests, which only fetch headers, not response bodies. This is much faster for status checks on pages with large HTML or images. If a server returns 405 (Method Not Allowed) for HEAD, the script automatically falls back to GET.
Some servers misbehave on HEAD requests (returning wrong status codes or rejecting them outright). The GET fallback handles that. Use --head when checking large lists of content-heavy pages; skip it when you need to verify the page actually loads.
Error types
| Status | Meaning | When it fires |
|---|---|---|
| OK | 2xx response | URL is working |
| CLIENT_ERROR | 4xx response | 404 Not Found, 403 Forbidden, etc. |
| SERVER_ERROR | 5xx response | 500 Internal Server Error, 503, etc. |
| TIMEOUT | Request timed out | Server too slow or unreachable |
| CONNECTION_ERROR | DNS or network failure | Host doesn’t exist, DNS resolution failed |
| SSL_ERROR | Certificate problem | Expired, self-signed, or mismatched cert |
| ERROR | Other request failure | Catch-all for unexpected issues |
SSL/TLS handling
Self-signed and expired certificates are common on internal server fleets. The script catches SSLError as a distinct error category so you can see at a glance which URLs have cert problems.
For internal hosts with self-signed certs, pass --insecure to disable TLS verification:
uv run url_checker.py internal-urls.txt --insecureSecurity warning
Never use --insecure on public-facing URLs. It disables certificate verification entirely, which means the connection is vulnerable to man-in-the-middle attacks. Only use it for internal or intranet hosts where you control the network.
Running it
Quick start
- Save the script as
url_checker.py - Create a
urls.txtwith one URL per line (#for comments):
# URLs to check
https://www.google.com
https://www.github.com
https://nonexistent-website-12345.com
bitdoze.com
example.com- Run it:
uv run url_checker.py urls.txtCommon flag combinations:
# Check with 20 workers and 15s timeout
uv run url_checker.py urls.txt --workers 20 --timeout 15
# Use HEAD requests, save broken URLs to a custom file
uv run url_checker.py urls.txt --head --output broken.txt
# Export all results to CSV
uv run url_checker.py urls.txt --csv results.csvLock the script for reproducibility
Lock the dependency resolution so every run uses the same package versions:
uv lock --script url_checker.py
# Creates url_checker.py.lockIf you run the script inside a repo that has a pyproject.toml, uv 0.12+ discovers the project from the script’s directory and may pull in unrelated project dependencies. Use --no-project to avoid that:
uv run --no-project url_checker.py urls.txt --workers 10Test it yourself
Here’s a deterministic test file using httpbin.org endpoints. Create this as urls.txt:
https://httpbin.org/status/200 # should be OK
https://httpbin.org/status/404 # should be CLIENT_ERROR (broken)
https://httpbin.org/status/500 # should be SERVER_ERROR (broken)
https://httpbin.org/delay/5 # should TIMEOUT with --timeout 3
https://nonexistent-website-12345.invalid # should be CONNECTION_ERROR
https://expired.badssl.com # should be SSL_ERRORRun it:
uv run url_checker.py urls.txt --timeout 3 --workers 4Expected results:
| URL | Expected status | Why |
|---|---|---|
| httpbin.org/status/200 | OK | 2xx response |
| httpbin.org/status/404 | CLIENT_ERROR | 404 is not 2xx |
| httpbin.org/status/500 | SERVER_ERROR | 500 is not 2xx |
| httpbin.org/delay/5 | TIMEOUT | 5s delay exceeds 3s timeout |
| nonexistent…invalid | CONNECTION_ERROR | DNS resolution fails |
| expired.badssl.com | SSL_ERROR | Expired certificate |
Verify before publishing
If you adapt this for your own blog or documentation, run the test file yourself first. The httpbin.org and badssl.com endpoints are public services, so verify they’re live at publish time. The httpstat.us endpoints used in older versions of this article were unreachable during research; httpbin.org is more reliable.
Reading the output
| Code | Meaning | What to do |
|---|---|---|
| 200 | OK | Nothing, link works |
| 301/302 | Redirect | Script follows redirects; check the final destination |
| 403 | Forbidden | May need auth headers or a realistic User-Agent |
| 404 | Not found | Fix or remove the link |
| 429 | Rate limited | Reduce --workers or add delays between requests |
| 500 | Server error | Contact site owner or retry later |
| TIMEOUT | Too slow | Increase --timeout or check connectivity |
| CONNECTION_ERROR | DNS/network | Check URL spelling and DNS resolution |
| SSL_ERROR | Certificate problem | Use --insecure for internal hosts only |
The script writes broken URLs to problematic_urls.txt (or whatever you pass with --output), grouped by error type with a timestamp header. It also exits with code 0 when all URLs are OK and code 1 when any are broken, which is what makes CI integration work.
If you need to load-test the endpoints you’re checking, take a look at oha for website performance testing as a companion tool.
Schedule it in CI with GitHub Actions
Here’s a complete workflow that runs the URL checker weekly and fails the build when broken links are found:
name: URL Check
on:
schedule:
- cron: "0 9 * * 1" # weekly, Monday 09:00 UTC
workflow_dispatch: # allow manual runs
jobs:
check:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@7d8f2266908b6c964bc9f6b357c94adc9594baf3 # v7.0.1
- name: Install uv
uses: astral-sh/[email protected]
with:
version: "0.12.6"
enable-cache: true
prune-cache: true
- name: Run URL check
run: uv run --no-project url_checker.py urls.txt --workers 20 --output problematic_urls.txtA few things to know:
- Pin immutable tags for setup-uv. Version 8.0.0 removed
@v8-style major/minor tags. They no longer resolve. Use@v10.0.1(immutable tag) or a full commit SHA. --no-projectin repos. Since uv 0.12.0,uv rundiscovers the project from the script’s directory. If your checker lives in a repo with apyproject.toml, pass--no-projectto avoid pulling in unrelated dependencies.prune-cache: true. Setup-uv v9.0.0 flipped the default tofalse, meaning the cache can grow and eat into your Actions cache quota (100 GB soft cap per repo). Setprune-cache: trueto keep it lean.workflow_dispatch. Lets you trigger the check manually from the Actions tab, useful after fixing a batch of broken links.- Exit code 1. The script exits 1 when broken links are found, so the workflow actually fails instead of silently passing.
Cost note
GitHub Actions free tier covers weekly checks of a few hundred URLs with no issues. If you prefer running from a VPS, a cron job on a Hetzner Cloud VPS works fine. The script is read-only GETs with zero blast radius. Schedule it with 0 9 * * 1 and redirect output to a log file.
VPS prices jumped across the board in 2026 — if you’re rethinking a rented box, see what changed and when a mini PC wins.
For always-on monitoring (not just weekly cron), consider a self-hosted Uptime Kuma instance or a Beszel and Uptime Kuma monitoring stack for broader server observability.
When Python is overkill: use lychee
If link checking is the entire job, no custom auth headers, no scoring, no database persistence, a dedicated tool beats a Python script. lychee is a Rust-based link checker: single static binary, async, per-host throttling built in, caching, and exit code 1 on broken links.
Python script
When to use: You need custom logic — auth headers, custom scoring, DB persistence, API-specific checks, or embedding the checker into a larger pipeline.
Install: Nothing — uv run handles it.
Run:
uv run url_checker.py urls.txt --head --workers 20lychee
When to use: Link checking IS the job. Drop-in binary, no code to maintain, built-in per-host throttling.
Install:
# macOS
brew install lychee
# Linux (binary)
curl -LO https://github.com/lycheeverse/lychee/releases/latest/download/lychee-x86_64-unknown-linux-gnu.tar.gz
tar xzf lychee-x86_64-unknown-linux-gnu.tar.gzRun:
lychee --verbose urls.txtAlso available as a GitHub Action (lycheeverse/lychee-action) and Docker image.
Tips and tricks
Send a realistic User-Agent
Some sites block requests that don’t look like browsers. The script sets a default User-Agent (url-checker/2.0 (bitdoze.com)). If you need a different one, change the session header near the top of the script:
SESSION.headers["User-Agent"] = "Mozilla/5.0 (compatible; MyBot/1.0)"Export results to CSV
Pass --csv results.csv to get a full CSV export of every URL with its status, status code, error type, and response time:
uv run url_checker.py urls.txt --csv results.csvThe CSV function uses Python’s built-in csv.DictWriter — no extra dependencies:
import csv
def save_to_csv(results, filename="results.csv"):
with open(filename, "w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=["url", "status", "status_code", "error_type", "response_time"])
writer.writeheader()
writer.writerows(results)Stay polite: per-host throttling
Ten to twenty threads hammering one domain will trigger bot protection on most sites. If your URL list targets a single domain, reduce --workers to 3-5. For mixed-domain lists, the ThreadPoolExecutor naturally spreads load across different hosts.
For heavy-duty per-host throttling, lychee has --host-concurrency and --host-request-interval built in.
Wrap up
This bulk URL checker gives you a single-file Python script that runs with uv run — no virtualenv, no pip, no setup. It checks hundreds of URLs concurrently, properly classifies 4xx and 5xx as broken (not “OK”), supports HEAD requests for speed, and exits with code 1 for CI integration. Lock it with uv lock --script for reproducible runs, schedule it in GitHub Actions or a VPS cron, and export results to CSV when you need them.
If you need to feed it a full site’s worth of URLs, you can export every WordPress post URL and title first. And for broader ops context — monitor your server and Docker resources to keep the infrastructure behind your sites healthy.
FAQ
How many URLs can this script check?
Hundreds comfortably with the default 10 workers. For 10,000+ URLs, increase --workers to 20-30, but be polite to target servers — too many concurrent requests to one domain triggers bot protection. For very large lists, consider batching by domain or using lychee’s async engine.
Can I check URLs that require authentication?
Yes. Add auth headers to the session near the top of the script:
SESSION.headers["Authorization"] = "Bearer YOUR_TOKEN"For cookie-based auth, set SESSION.cookies. The session is shared across all threads, so every request carries the same headers.
Why not just use curl in a loop?
You can, but you lose concurrency, error classification, and progress reporting. A for url in $(cat urls.txt); do curl ...; done loop checks one URL at a time and gives you raw HTTP codes with no grouping or output file. This script wraps the same HTTP logic with parallel execution, structured output, and CI-friendly exit codes.
Does this work on Windows?
Yes. uv runs on Windows, macOS, and Linux. The script uses only Python stdlib plus requests — no platform-specific code. Install uv with powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex" on Windows.
How do I check links across my whole site?
Export your sitemap or URL list first. For WordPress sites, export every WordPress post URL and title to get a complete list, then feed it to this script. For static sites, parse your sitemap.xml or use a crawler to generate the URL list.


