8.0 KiB
Guardian Daily Newscast Pipeline
Turns The Guardian's daily US news headlines into a spoken MP3 newscast.
Three scripts run in sequence: scrape the articles, generate teleprompter scripts via a local LLM, convert scripts to audio.
Prerequisites
External tools you must have installed before running:
- Ollama — with the
gemma4:e2bmodel pulled (ollama pull gemma4:e2b). Used in stage 2 to generate newscast scripts. - ffmpeg — used in stage 3 to convert WAV to MP3 (
brew install ffmpegon macOS,apt install ffmpegon Ubuntu).
Quick Start
# 1. Install Python dependencies
pip install -r requirements.txt
# 2. Run the full pipeline (stages 1, 2, 3 in order)
python main.py && python reader.py && python tts_generator.py
# 3. Find your newscast files in data/audio/*.mp3
ls data/audio/
Installation
pip install -r requirements.txt
Dependencies: httpx, beautifulsoup4, pydantic, loguru, kokoro, soundfile, torch.
The Pipeline
The project has three distinct stages that must be run in order:
[Stage 1] main.py (Guardian Scraper)
|
v
data/output.json
|
v
[Stage 2] reader.py (LLM Script Generator)
|
v
data/scripts.json
|
v
[Stage 3] tts_generator.py (TTS Audio Generator)
|
v
data/audio/*.mp3
Stage 1: Guardian Scraper (main.py)
- Fetches today's Guardian US news home page (URL built from current date).
- Extracts all article links, then fetches each article's body text concurrently (semaphore-limited to 5).
- Output:
data/output.json— a JSON object with anarticlesarray (each entry:headline,url,body,scraped_at).
Stage 2: LLM Script Generator (reader.py)
- Reads
data/output.json. - Sends each article body to a local Ollama instance (
http://127.0.0.1:11434/api/generate, modelgemma4:e2b). - System prompt instructs the model to produce clean teleprompter-ready spoken text (no stage directions, no speaker labels).
- Retries failed requests up to 3 times (exponential backoff: 10s, 20s). 2s delay between each request to avoid overloading Ollama.
- Output:
data/scripts.json— a JSON array of objects (each:headline,script).
Stage 3: TTS Audio Generator (tts_generator.py)
- Reads
data/scripts.json. - Uses Kokoro TTS (local, voice
af_heart, 24kHz) to generate WAV audio. - Converts WAV to MP3 via ffmpeg (
-q:a 2). - Cleans up intermediate WAV files after conversion.
- Output:
data/audio/<sanitized_headline>.mp3files — one MP3 per article.
Usage
Run all three stages in order:
python main.py
python reader.py
python tts_generator.py
Or chain them with &&:
python main.py && python reader.py && python tts_generator.py
Each stage is independent and can be re-run without re-running earlier stages (e.g., re-run stage 3 if audio quality needs to change, without re-scraping).
Project Structure
.
├── main.py # Stage 1: Guardian scraper
├── reader.py # Stage 2: LLM script generation
├── tts_generator.py # Stage 3: TTS audio generation
├── pipeline.py # Pipeline orchestrator (Docker)
├── requirements.txt # Python dependencies
├── README.md # This file
├── CLAUDE.md # Developer guide
├── Dockerfile # Docker container build
├── docker-compose.yml # Docker Compose (optional Ollama sidecar)
├── .env.example # Environment variable defaults
├── server/
│ ├── __init__.py
│ └── app.py # Web server (MP3 index + file serving)
├── scraper/
│ ├── __init__.py
│ ├── client.py # GuardianClient (HTTP + concurrency)
│ ├── parser.py # GuardianParser (HTML extraction)
│ ├── models.py # Pydantic models (Article, ScrapeResult)
│ └── utils.py # Disk I/O helpers
├── data/
│ ├── output.json # Stage 1 output (scraped articles)
│ ├── scripts.json # Stage 2 output (generated scripts)
│ └── audio/ # Stage 3 output (MP3 files)
│ ├── article_1.mp3
│ ├── article_2.mp3
│ └── ...
└── .claude/
└── settings.local.json
Dependencies
Python (via pip install -r requirements.txt)
| Package | Purpose |
|---|---|
| httpx | Async HTTP client for scraping |
| beautifulsoup4 | HTML parsing |
| pydantic | Data validation models |
| loguru | Logging |
| kokoro | Local TTS engine |
| soundfile | WAV audio I/O |
| torch | Required by kokoro |
External (not installed via pip)
| Tool | Requirement |
|---|---|
| Ollama | Running locally, with gemma4:e2b |
| ffmpeg | Available on PATH |
Output Files
| File | Description |
|---|---|
data/output.json |
JSON object containing scraped articles with headline, url, body, and scraped_at fields. |
data/scripts.json |
JSON array of objects with headline and script fields. |
data/audio/*.mp3 |
One MP3 file per article, named from the sanitized headline. |
Docker Deployment
Run the pipeline as a self-contained Docker container with a built-in web UI for browsing and downloading newscast files.
Quick Start
# 1. Build the image
docker build -t guardian-newscast .
# 2. Run (web UI on port 8080, data persisted in ./data/)
docker run -d \
--name guardian-newscast \
-p 8080:8080 \
-v ./data:/app/data \
guardian-newscast
# 3. Open the web UI
open http://localhost:8080
Volume Mapping
| Mount Point | Purpose |
|---|---|
./data:/app/data |
Required — stores all pipeline outputs (output.json, scripts.json, audio/*.mp3). Host path is your choice. |
Environment Variables
| Variable | Default | Purpose |
|---|---|---|
OLLAMA_URL |
http://host.docker.internal:11434/api/generate |
Ollama API endpoint |
MODEL_NAME |
gemma4:e2b |
Ollama model for script generation |
PORT |
8080 |
Web server HTTP port |
POLL_INTERVAL |
21600 (6h) |
Seconds between pipeline runs |
DATA_DIR |
/app/data |
Data directory inside container |
TTS_VOICE |
af_heart |
Kokoro TTS voice name |
Override via .env file (copy .env.example), docker run -e, or docker-compose.
Web Server
The container serves a dark-themed page at / listing all MP3 files with:
- Inline audio playback
- Direct download links
- Auto-refresh every 5 minutes
Access at http://<host>:<PORT> (default http://localhost:8080).
Publishing
docker build -t guardian-newscast .
docker tag guardian-newscast truenas.local:30095/guardian-newscast:latest
docker push truenas.local:30095/guardian-newscast:latest
Cross-platform builds
For Intel (x86_64):
docker buildx build --platform linux/amd64 -t guardian-newscast:amd64 .
For ARM (aarch64 / Apple Silicon):
docker buildx build --platform linux/arm64 -t guardian-newscast:arm64 .
To build and push both variants as a single multi-arch manifest:
docker buildx build --platform linux/amd64,linux/arm64 -t truenas.local:30095/guardian-newscast:latest --push .
Docker Compose
# Minimal (pipeline only, Ollama on host)
docker compose up -d
# Full (pipeline + Ollama sidecar)
docker compose --profile full up -d
GPU Acceleration
The default Dockerfile uses python:3.11-slim (Debian), which has glibc — fully compatible with NVIDIA CUDA.
- Build with CUDA-enabled torch:
(Add a
docker build --target gpu -t guardian-newscast:gpu .FROM python:3.11-slim AS gpustage withpip install torch --index-url https://download.pytorch.org/whl/cu121) - Run with
--gpus all. - Minimum 4 GB VRAM recommended for Ollama + Kokoro in the same container.