Overview & Version 4
BaseCore.fetch(), BaseMedia.load(api=..., html=...), mutable ScrapeResult.video/is_success contract, and boolean retry callbacks were removed. Use the explicit interfaces documented below.
eaf_base_api is the shared engine used by all site-specific API packages. Version 4 separates each responsibility into an explicit, typed interface:
- BaseCore — HTTP sessions, request retries, response decoding, caching, proxy/interface binding, HLS inspection, and downloads
- BaseMedia + media_field — source-aware, atomic lazy loading with deterministic source precedence
- Helper + IteratorConfig — bounded page/item scheduling, ordering, retries, error handlers, and load selection
- ScrapeStream + ScrapeResult — deterministic stream cleanup and immutable success/failure results
- CacheBackend — replaceable storage contract; the built-in cache uses byte limits and TTL expiry
Installation
pip install eaf_base_api
# Install optional HLS parsing/remux dependencies as well
pip install "eaf_base_api[hls]"
Version 4 requires Python 3.12 or newer. The package ships a py.typed marker, so type checkers can consume its inline annotations.
RuntimeConfig
Create a RuntimeConfig for each independently configured BaseCore. A process-wide config instance remains available as the default, but a dedicated instance avoids unrelated clients changing one another.
| Attribute | Type | Default | Description |
|---|---|---|---|
response_cache_size_bytes | int | 32 MiB | Maximum encoded size of cached text responses; set to 0 to disable |
response_cache_ttl | float | 300.0 | Text-response lifetime in seconds |
segment_cache_size_bytes | int | 8 MiB | Maximum encoded size of cached HLS segment URL lists |
segment_cache_ttl | float | 300.0 | HLS segment-list lifetime in seconds |
request_attempts | int | 4 | Total request attempts, including the first call |
request_retry_initial_delay | float | 0.5 | Initial exponential retry delay in seconds |
request_retry_max_delay | float | 30.0 | Maximum exponential retry delay |
request_multiplier | float | 2.0 | Exponential base for BaseCore request backoff and multiplier for derived page/item retry policies |
request_retry_jitter | float | 0.5 | Maximum random jitter added to retry delays |
request_delay | int | 0 | Minimum delay between requests made by this core |
timeout | int | 20 | Default request timeout in seconds |
max_bandwidth_mb | float | None | None | Aggregate download receive limit in MB/s |
proxy | str | None | None | One HTTP, HTTPS, or SOCKS proxy URL |
proxy_auth | str | None | None | Proxy credentials as "username:password" |
interface | str | None | None | Local interface IP address to bind |
ip_resolve | int | None | None | Optional IP resolution: None (dual-stack default), 1 (IPv4 only), or 2 (IPv6 only). Also respects CURL_IPRESOLVE environment variable. |
http_version | str | "v2" | "v1", "v2", or "v3" |
dns_over_https | str | None | None | DNS-over-HTTPS endpoint |
impersonation | str | "chrome" | curl_cffi browser impersonation profile |
custom_ja3 | str | None | None | Advanced custom TLS JA3 fingerprint |
verify_ssl | bool | True | Verify TLS certificates |
trust_env | bool | False | Use proxy and CA settings from the environment |
cookies | dict[str, str] | None | None | Initial cookie mapping applied when each new HTTP session is created |
locale | str | "en-US,en;q=0.9" | Default Accept-Language; changing it can affect site parsers |
max_workers_download | int | 20 | HLS download worker count for this core |
videos_concurrency | int | 5 | Fallback item concurrency for resolved iterator configs |
pages_concurrency | int | 2 | Fallback page concurrency for resolved iterator configs |
When changes take effect
Cache limits/TTLs and the initial locale header are captured when BaseCore is constructed. Proxy, interface, ip_resolve, TLS, HTTP, DoH, impersonation, cookies, and bandwidth options are captured whenever its session is created. Request budgets, retry delays/multiplier, and request delay are read for each request; HLS timeout/workers are read when a download begins; and omitted IteratorConfig values are resolved whenever a new stream is created.
initialize_session() is idempotent: it creates a session only while core.session is None. To apply changed session-bound settings such as proxy, cookies, TLS, or interface, call await core.close(); the next request, context-manager entry, or explicit initialize_session() creates a fresh session from the current configuration.
Client Integration & Lifecycle
from base_api import BaseCore
from base_api.modules.config import RuntimeConfig
from xvideos_api import Client
runtime = RuntimeConfig()
runtime.timeout = 60
runtime.request_attempts = 5
runtime.request_multiplier = 2.0
runtime.proxy = "socks5://127.0.0.1:9050"
runtime.interface = None # Or a local interface IP
runtime.cookies = {"session": "value"}
core = BaseCore(configuration=runtime)
client = Client(core=core)
try:
video = await client.get_video("https://www.xvideos.com/video...")
finally:
await core.close()
BaseCore starts without an HTTP session. Context-manager entry, the first request, or an explicit initialize_session() creates it lazily; further initialization calls are no-ops while that session remains live. Prefer async with BaseCore(configuration=runtime) as core: for direct use. await core.close() closes and clears the current session, and a later request safely recreates it from the core's current configuration.
Explicit Request API
Version 4 replaces the flag-driven fetch() method with one method per representation:
| Method | Returns | Cache behavior |
|---|---|---|
await core.request(url, ...) | curl_cffi.Response | Never uses the text cache |
await core.fetch_text(url, ...) | str | Successful GET text can use the configured cache |
await core.fetch_bytes(url, ...) | bytes | Never uses the text cache |
from base_api import BaseCore, CachePolicy
async with BaseCore() as core:
response = await core.request(url)
text = await core.fetch_text(url)
data = await core.fetch_bytes(download_url)
refreshed = await core.fetch_text(
url,
cache_policy=CachePolicy.REFRESH,
)
uncached = await core.fetch_text(
url,
cache_policy=CachePolicy.BYPASS,
)
CachePolicy.USE reads and writes, REFRESH skips the read and replaces the cached value, and BYPASS neither reads nor writes. All three request methods accept timeout, cookies, redirects, data, HTTP method, headers, JSON, parameters, and retry_non_idempotent.
BaseCore Request Retries
RuntimeConfig.request_attempts includes the first request. Backoff uses request_retry_initial_delay, request_multiplier as Tenacity's exponential base, request_retry_max_delay, and request_retry_jitter. Idempotent methods (GET, HEAD, PUT, DELETE, OPTIONS, and TRACE) retry network failures plus HTTP 408, 425, 429, and 5xx responses. A 429 response honors Retry-After when available.
- POST/PATCH and other non-idempotent operations are attempted once unless
retry_non_idempotent=Trueis explicitly safe. - Non-retryable HTTP failures—including 401/403, 404, 410, and 4xx responses other than 408, 425, and 429—are terminal.
- Exhaustion raises
RequestRetriesExhaustedwith.url,.attempts, and.last_error.
RuntimeConfig controls individual HTTP requests. RetryPolicy controls complete page extraction or media-item loading inside Helper. When those policies are derived, both layers use RuntimeConfig.request_multiplier; avoid multiplying both budgets unnecessarily.
Proxy, Interface & IP Resolution
from base_api.modules.config import RuntimeConfig
runtime = RuntimeConfig()
runtime.proxy = "socks5://127.0.0.1:9050"
runtime.proxy_auth = "username:password"
runtime.interface = "192.0.2.10"
runtime.ip_resolve = 1 # 1 for IPv4, 2 for IPv6, None for dual-stack
# Only disable verification for a proxy you control and trust.
runtime.verify_ssl = False
The old proxies mapping was replaced by the singular proxy URL. interface is passed to curl_cffi.AsyncSession and must be a local interface IP address. ip_resolve allows restricting DNS resolution to IPv4 (1) or IPv6 (2) when dual-stack connectivity issues arise; it also respects the CURL_IPRESOLVE environment variable (e.g. CURL_IPRESOLVE=1 or 4 for IPv4, 2 or 6 for IPv6).
Caching
The built-in Cache uses separate, thread-safe TTL caches for text responses and HLS segment lists. Limits are measured in UTF-8 bytes, not item counts. Request keys include method, URL, redirects, parameters, body, headers, and cookies; sensitive values are fingerprinted rather than stored in plaintext keys.
- Only successful GET text responses are stored.
- Concurrent cache misses for the same request share one in-flight network operation.
core.cache.clear()clears both built-in caches.- Pass a custom
CacheBackendtoBaseCore(cache=...)to replace storage.
Source-aware BaseMedia
Remote fields are declared with media_field(). The first source has highest precedence. A loader is async, returns a complete mapping, and never mutates the object directly; results are validated and committed atomically.
from dataclasses import dataclass
from typing import ClassVar
from base_api import BaseMedia, media_field
@dataclass(kw_only=True, slots=True)
class Video(BaseMedia):
title: str | None = media_field("html", "api")
duration: int | None = media_field("api")
loader_methods: ClassVar[dict[str, str]] = {
"html": "_load_html",
"api": "_load_api",
}
async def _load_html(self) -> dict[str, object]:
data = await fetch_html(self.url)
return {"title": data.get("title")}
async def _load_api(self) -> dict[str, object]:
data = await fetch_api(self.url)
return {"title": data.get("title"), "duration": data.get("duration")}
| Operation | Meaning |
|---|---|
await media.load_sources("html", "api") | Load named sources concurrently; repeated calls are idempotent |
await media.load_fields("title") | Select the smallest useful source set for unresolved fields |
await media.get_field("title") | Load one field if necessary and return it |
media.loaded_sources | Immutable set of successful source names |
media.source_state("html") | NOT_LOADED, LOADING, LOADED, or FAILED |
media.source_errors | Copy of the latest failure per source |
media.is_field_loaded(name) | True even when a loader deliberately returned None |
media.unloaded_fields() | Names still holding the private unloaded marker |
media.to_dict() | Serialize loaded public fields without triggering lazy-field errors |
Direct access to an unresolved field raises DataNotLoadedError. Use retry_failed=False with a load method when a previously failed source should not be attempted again.
IteratorConfig
Site-specific iterator methods now accept one IteratorConfig instead of many concurrency, ordering, loading, and callback parameters. Values left as None are resolved from the active core's RuntimeConfig.
| Attribute | Default | Description |
|---|---|---|
max_page_concurrency | RuntimeConfig.pages_concurrency | Concurrent page operations |
max_item_concurrency | RuntimeConfig.videos_concurrency | Concurrent media-item operations |
max_pending_items | 4 × item concurrency | Backpressure limit for extracted items awaiting work |
extract_in_thread | True | Run the synchronous extractor and its iteration in a worker thread |
order | ResultOrder.COMPLETION | Yield fastest results first or restore original page/item order |
page_error_mode | ErrorMode.YIELD | Terminal page failure behavior |
item_error_mode | ErrorMode.YIELD | Terminal item failure behavior |
page_retry | derived from RuntimeConfig | Bounded retry policy for a complete page operation |
item_retry | derived from RuntimeConfig | Bounded retry policy for construction/loading of one item |
page_error_handler | None | Optional sync/async handler receiving ScrapeErrorContext |
item_error_handler | None | Optional sync/async handler receiving ScrapeErrorContext |
load_specific_fields | () | Fields each constructed media item must load |
load_specific_sources | () | Sources each constructed media item must load before fields |
IteratorConfig replaces that API method's default config. Include the method's required load_specific_sources or load_specific_fields; most site packages use ("html",), while some dual-source APIs use both "api" and "html".
RetryPolicy & Custom Error Handling
RetryPolicy.max_attempts includes the first attempt. Without a custom handler, eligible exception classes are selected with retry_for; the delay is bounded exponential backoff plus uniformly random jitter.
RetryPolicy() performs one stage attempt. When an IteratorConfig leaves page_retry or item_retry unset, resolve() instead creates both policies from the active RuntimeConfig; the shipped defaults are four stage attempts, base_delay=0.5, multiplier=2.0, max_delay=30.0, and jitter=0.5. A stage attempt may call BaseCore, whose separate four-attempt request budget can therefore multiply the number of HTTP calls. Set both layers deliberately for your workload.
from base_api import (
ErrorAction,
ErrorMode,
MediaLoadError,
MediaLoadErrors,
ResultOrder,
RetryPolicy,
ScrapeErrorContext,
)
from base_api.modules.config import IteratorConfig
from base_api.modules.errors import ResourceGone
def resource_is_gone(error: BaseException) -> bool:
if isinstance(error, ResourceGone):
return True
if isinstance(error, MediaLoadError):
return resource_is_gone(error.original_error)
if isinstance(error, MediaLoadErrors):
return any(resource_is_gone(item) for item in error.errors)
return False
def handle_page_error(context: ScrapeErrorContext) -> ErrorAction:
print("page", context.url, context.attempt, context.error)
return ErrorAction.RETRY
async def handle_item_error(context: ScrapeErrorContext) -> ErrorAction:
print("item", context.url, context.attempt, context.error)
if resource_is_gone(context.error):
return ErrorAction.SKIP
return ErrorAction.RETRY
iterator_config = IteratorConfig(
max_page_concurrency=2,
max_item_concurrency=8,
order=ResultOrder.ORIGINAL,
item_retry=RetryPolicy(
max_attempts=3,
base_delay=0.5,
multiplier=2.0,
max_delay=8.0,
jitter=0.25,
retry_for=(Exception,),
),
page_error_handler=handle_page_error,
item_error_handler=handle_item_error,
item_error_mode=ErrorMode.YIELD,
page_error_mode=ErrorMode.SKIP,
load_specific_sources=("html",),
)
Page and item handlers are independent; each receives only failures from its own stage, and either field may be left unset to use automatic policy for that stage. A sync or async handler runs on every failed stage attempt and may return RETRY, RAISE, YIELD, or SKIP. An explicit RETRY can override retry_for, but it cannot exceed the policy's hard limit; on the final attempt it falls back to the configured error mode. Page failures containing a nested HTTP 404 remain terminal. A handler that raises or returns another value produces a fatal ErrorHandlerError.
ScrapeStream & ScrapeResult
Helper.iterator() returns a lazily started ScrapeStream async context manager. For convenient direct iteration without manually managing context lifecycles, scrape_stream() provides an AsyncGenerator[ScrapeResult[T], None] that handles internal stream lifecycle automatically.
# Direct iteration via scrape_stream() async generator
from base_api import scrape_stream
async for result in scrape_stream(
target_page_urls=page_urls,
item_extractor=extractor,
iterator_config=iterator_config,
core=core,
):
if not result.succeeded:
print(result.stage, result.url, result.error)
continue
media = result.unwrap() # Returns T or raises result.error
# Alternatively, using Helper.iterator() with an async context manager
stream = helper.iterator(
target_page_urls=page_urls,
item_extractor=extractor,
iterator_config=iterator_config,
)
async with stream:
async for result in stream:
if not result.succeeded:
print(result.stage, result.url, result.error)
continue
media = result.unwrap()
scrape_stream() is an async generator: use async for result in scrape_stream(...): directly without an async with block. In contrast, Helper.iterator() returns a ScrapeStream context manager which supports async with stream: to guarantee background worker cancellation if the loop exits early.
| ScrapeResult field | Description |
|---|---|
stage | ScrapeStage.PAGE or ScrapeStage.ITEM |
url | Page or item URL associated with this outcome |
page_index / item_index | Stable original ordering coordinates |
attempts | Number of attempts consumed |
item | Loaded media on success, otherwise None |
error | Typed PageFetchError/ItemFetchError on yielded failure |
succeeded | True exactly when the result contains an item |
unwrap() | Return the item or raise the stored typed scrape error |
Download Configurations
Shared fields
DownloadConfigHLS and DownloadConfigRAW share quality, path, callback, no_title, and stop_event. Quality accepts a supported integer height or "best", "half", "worst", and common "720p"-style values. Version 4.1 normalizes landscape and portrait variants by their shorter dimension, chooses the closest available tier for an explicit numeric request (preferring the higher tier on a tie), and uses bandwidth and frame rate as fallbacks when a master playlist omits resolution metadata.
HLS inspection
await core.list_available_qualities(m3u8_url) returns sorted, unique integer tiers from a master playlist. await core.get_m3u8_by_quality(m3u8_url, quality) returns the selected media-playlist URL. Both methods accept a playlist URL, inline master text, or a callable/awaitable resolving to either form. Relative variant paths can be resolved only when the master was fetched from a URL; an inline master must contain absolute variant URLs.
DownloadConfigHLS
| Field | Default | Description |
|---|---|---|
m3u8_base_url | None | Master playlist URL, awaitable, or callable used by BaseCore.download |
remux | False | Remux concatenated transport stream to MP4 |
start_segment | 0 | First segment index |
segment_state_path | None | JSON resume-state file; Path values are serialized safely in 4.1+ |
segment_dir | None | Directory for downloaded segments |
return_report | False | Return a DownloadReport instead of only a boolean |
cleanup_on_stop | True | Remove temporary state when cancelled |
keep_segment_dir | False | Retain segment files after completion |
callback_remux | None | Progress callback for remuxing |
ios_support | False | Enable the iOS-compatible remux path |
BaseCore.download() reads RuntimeConfig.timeout and RuntimeConfig.max_workers_download from that core when dispatching each HLS download. Dedicated cores therefore keep their HLS timeout and worker settings isolated from the process-wide default configuration. HLS cancellation uses an asyncio.Event: setting it cancels pending segment tasks promptly, and progress/remux callbacks are synchronized across worker threads. Version 4.1.1 also normalizes discontinuous packet timestamps while remuxing so valid HLS discontinuities do not abort MP4 output.
remux=True) requires the PyAV library. Install it with pip install "eaf_base_api[hls]" or pip install av.
DownloadConfigRAW
Direct-file downloads add allow_multipart=True, max_workers=5, read_timeout=120.0, chunk_size=1024, and max_retries=5. These downloader retries are separate from RuntimeConfig.request_attempts.
Provider Helpers (4.2.0)
base_api.modules.provider centralizes common request and download handling across all 15 site packages, eliminating boilerplate while enforcing uniform exception translation and contextual logging:
| Function | Signature | Description |
|---|---|---|
fetch_content | async (core, url, *, logger, owner=None, get_json=False, error_types=errors, not_found_error=None, access_denied_error=None) -> str | dict | Fetches text or JSON and translates transport failures (HTTP 404, bot blocks, proxy errors, network errors) into provider exceptions, chaining the cause. Attaches class_name and url context. |
download_errors | (error_type=errors.DownloadFailed) -> decorator | Decorator giving media download methods uniform logging, contextual tracing, and cancellation semantics. Raises DownloadFailed (with URL, class_name, and API) on failure, while propagating DownloadCancelled and asyncio.CancelledError cleanly. |
prepare_download_config | (configuration, title: str | None) -> configuration | Shallow-copies download configuration and appends sanitized title with .mp4 when no_title=False, keeping callbacks and stop events shared. |
download_hls | async (media, configuration) -> bool | DownloadReport | Loads title and m3u8_base_url, prepares the configuration, and delegates to media.core.download(configuration=config). |
from base_api.modules.provider import fetch_content, download_errors, download_hls
from base_api.modules.logger import get_logger
from my_api.modules.errors import DownloadFailed
logger = get_logger(__name__)
async def get_html_content(core, url, *, owner=None):
return await fetch_content(core, url, logger=logger, owner=owner)
@download_errors(DownloadFailed)
async def download(self, configuration):
return await download_hls(self, configuration)
Shared Utilities
base_api exports standardized parsing, sanitization, and helper functions used across scrapers:
| Function | Signature | Description |
|---|---|---|
get_text_safe | (node: Any, selector: str | None = None) -> str | None | Safely extracts stripped text from a node or CSS selector match, returning None if missing or empty. |
get_attr_safe | (node: Any, attr: str, selector: str | None = None) -> str | None | Safely extracts an HTML attribute string from a node or CSS selector match. |
parse_duration | (text: str | None) -> int | None | Parses duration strings (e.g. "12:34", "1:23:45", "45s", "12m") into integer seconds. |
parse_count | (text: str | int | float | None) -> int | None | Parses view and like counts with human-readable suffixes (e.g. "1.5M", "300K", "12,450") into integers. |
build_m3u8_master | (streams: list[dict[str, Any]] | dict[str, str]) -> str | Constructs a standard HLS master playlist (#EXTM3U) with resolution and bandwidth headers from a list or dictionary of stream variants. |
extract_video_grid | (parser: Any, base_url: str, logger: Any) -> list[dict] | Extracts standardized video attributes (title, url, thumbnail, duration, author, preview) from search/listing grids. Used by Tube8 and Thumbzilla. |
str_to_bool | (value: Any) -> bool | Coerces common string boolean representations ("true", "1", "yes") to boolean. |
is_resource_gone | (exc: BaseException) -> bool | Checks if an exception or nested cause represents an HTTP 404, 410, or ResourceGone error. |
contains_resource_gone | (exc: BaseException) -> bool | Recursively inspects aggregate exceptions (e.g. MediaLoadErrors) to check if any nested failure represents a gone resource. |
Logging & Cleanup
Configure logging once at application startup to capture provider and base API records in the same console or file output:
import logging
from base_api import configure_app_logging, DownloadFailed
configure_app_logging(log_file="api.log", level=logging.INFO)
With no logger_name, configure_app_logging() configures the root logger. Omit log_file for console output only. Log files append by default; overwrite_file=True truncates the selected file. The default format includes timestamp, logger name, level, filename, line number, function, and message. Logged exceptions include the original traceback and chained causes.
Contextual Logging (4.2.0)
Every library logger is wrapped in a ContextLogger adapter obtained via get_logger(name, owner=None). When operations run inside a log_context(owner, url) block, messages are automatically prefixed with [class=<class_name> url=<url>], and the context variables are injected into LogRecord.class_name and LogRecord.url. Because context is stored in ContextVar instances, concurrent downloads and background tasks offloaded via asyncio.to_thread keep their log context completely separate.
ERROR xvideos_api.api [class=xvideos_api.api.Video url=https://example.test/video/123] Download failed: disk full
Traceback (most recent call last):
...
OSError: disk full
Creating a client or core does not attach handlers or change your application's logging configuration. Without application setup, Python's default logging handler prints warnings and errors to stderr.
Provider CLI entry points configure INFO-level console logging automatically and log caught per-URL failures with tracebacks. Library imports do not configure the root logger; importing Tube8 or Thumbzilla no longer enables root DEBUG logging.
Logging a caught or stored exception
Use logger.exception() inside an except block. The APIs already log failures at their handling boundaries; add application logging when more context is needed. Outside the original handler, pass the stored exception explicitly:
logger = logging.getLogger(__name__)
try:
await video.download(configuration=config)
except DownloadFailed as error:
print(error.api, error.class_name, error.url)
# The failure and its full traceback have already been logged.
except Exception:
logger.exception("Application download failed for %s", video.url)
raise
if not result.succeeded:
error = result.error
logger.error(
"Scrape failed for %s: %s", result.url, error,
exc_info=(type(error), error, error.__traceback__),
)
exc_info=True alone cannot recover a traceback from a stored ScrapeResult. Ordinary download failures now raise DownloadFailed instead of returning False. Explicit cancellation via stop_event raises DownloadCancelled, and async task cancellation propagates asyncio.CancelledError without error logs. With DownloadConfigHLS(return_report=True), incomplete or cancelled downloads still return their explicit DownloadReport.
Core logging and cleanup
core.enable_logging() configures the core logger specifically. Use root logging above when the file should also receive provider and media-loader records. Optional remote logging uses log_ip/log_port on the core or http_ip/http_port on configure_app_logging().
import logging
core.enable_logging(level=logging.DEBUG)
core.enable_logging(log_file="api.log", level=logging.INFO)
core.enable_logging(
log_ip="192.168.1.100",
log_port=8080,
level=logging.DEBUG,
)
# Always release the curl_cffi connection pool.
await core.close()
Migration from 3.x
| 3.x | 4.0 replacement |
|---|---|
fetch(url) | fetch_text(url) |
fetch(..., get_bytes=True) | fetch_bytes(...) |
fetch(..., get_response=True) | request(...) |
fetch(..., save_cache=False) | fetch_text(..., cache_policy=CachePolicy.BYPASS) |
max_cache_items | response_cache_size_bytes + segment_cache_size_bytes |
max_retries | request_attempts |
proxies | proxy |
load(api=True, html=False) | load_sources("api") or load_fields(...) |
on_error_hint returning bool | ErrorHandler returning ErrorAction |
keep_original_order=True | IteratorConfig(order=ResultOrder.ORIGINAL) |
max_video_concurrency | IteratorConfig(max_item_concurrency=...) |
result.is_success | result.succeeded |
result.video | result.unwrap() or checked result.item |
Error Reference
Errors live in base_api.modules.errors; frequently used version 4 errors are also exported from base_api.
The common provider errors derive from ScraperException, which now derives from BaseScraperError. Catch common failures from base_api: DownloadFailed, NetworkError, NotFound, ProxyError, BotDetection, VideoUnavailable, LoginFailed, and RegionBlocked. Existing provider imports remain valid. Pornhub retains its own classes: its common errors subclass both the corresponding shared type and PornhubAPIError, preserving both ways of catching them.
Translated exceptions carry rich diagnostic attributes: error.url, error.class_name, and error.api. Provider request wrappers chain translated exceptions with raise ... from original_error, preserving lower-level causes in __cause__.
Download operations now raise DownloadFailed on failure instead of returning False. Download wrappers cover preparation failures (metadata loading, quality selection, output path setup); DownloadFailed includes the video URL and retains the original error in __cause__. An explicit DownloadCancelled or asyncio.CancelledError propagates without being wrapped as DownloadFailed and without error logs. With DownloadConfigHLS(return_report=True), incomplete or cancelled downloads still return their explicit DownloadReport.
MediaLoadError; inspect its original_error for a site package's NotFound, RegionBlocked, or similar exception. If several requested sources fail together, MediaLoadErrors.errors contains each failure. Helper item failures add one more typed ItemFetchError layer whose original_error is the media-load error.
| Family | Exceptions and meaning |
|---|---|
| Shared provider errors | BaseScraperError, ScraperException, NotFound, NetworkError, BotDetection, ProxyError, UnknownNetworkError, DownloadFailed, VideoUnavailable |
| HTTP/network | NetworkRequestError, HTTPStatusError, RateLimitError, RequestRetriesExhausted, ResourceGone, AccessDeniedError, InvalidProxy, ProxySSLError |
| Media fields | UnknownMediaFieldError, FieldNotLoadableError, DataNotLoadedError |
| Media loaders | LoaderConfigurationError, LoaderContractError, MediaLoadError, MediaLoadErrors |
| Scrape operations | PageFetchError, ItemFetchError, ErrorHandlerError |
| Downloads/playlists | DownloadCancelled, SegmentError, PlaylistExtractionError, StateLoadError, MaxRetriesExceeded |
| Bot challenges | BotProtectionDetected, ChallengeRegexError, ChallengeMathError, SecurityAbort |
Changelog
| Date / Version | Changes since 3.3.3 |
|---|---|
| 2026-09-15 — 4.2.0 | Centralized provider request and download helpers in base_api.modules.provider (fetch_content, download_errors, prepare_download_config, download_hls). Added contextual logging via ContextLogger, log_context, class_name, and get_logger with [class=... url=...] formatting. Structured exception reporting with error.url, error.class_name, and error.api attributes, and strict exception chaining via __cause__. Standardized download failure semantics: download() raises DownloadFailed on failure instead of returning False; cancellation via stop_event raises DownloadCancelled without error logs. Added extract_video_grid() to base_api.modules.static_functions for shared search result extraction. |
| 2026-09-07 — 4.1.2 | Added scrape_stream() async generator helper, optional dual-stack/IPv4/IPv6 address resolution via RuntimeConfig.ip_resolve and CURL_IPRESOLVE environment variable, standardized scraper utilities (get_text_safe, get_attr_safe, parse_duration, parse_count, build_m3u8_master, str_to_bool), and PyAV remuxing dependencies. |
| 2026-08-19 — 4.1.1 | Synchronized download/remux progress callbacks across worker threads, normalized discontinuous HLS packet timestamps during remuxing, and added regressions for stage-specific error handlers and per-core runtime settings. Commit a062695. |
| 2026-08-16 — 4.1.0 | Made HLS cancellation stop pending segment work promptly and corrected JSON resume-state serialization, including configured Path values. Commit acdb967. |
| 2026-08-14 — 4.0.1 | Reworked HLS quality discovery/selection for canonical landscape and portrait tiers, inline/callable playlists, and bandwidth fallbacks; fixed independent item error handlers, idempotent session initialization, configured session cookies and retry multipliers, and removed hard-coded transport headers. Commits 2fd3a55–1eb9105. |
| 2026-08-11 — 4.0.0 | Added the py.typed marker; finalized runtime-resolved IteratorConfig values; expanded typed exports; improved page/item exception logging and terminal error context. Commits 13b4105–184df57. |
| 2026-08-08 — 4.0.0 | Centralized page/item concurrency, ordering, loading, retry policies, and handlers in IteratorConfig; resolved omitted policies from RuntimeConfig; expanded inline API documentation. Commits 2705faa–baccdfd. |
| 2026-08-07 — 4.0.0 | Breaking v4 redesign: explicit request methods and cache policies; byte-bounded TTL caches and single-flight text fetches; source-aware atomic media loading; bounded dynamic Helper scheduling; immutable ScrapeResult; owned ScrapeStream; structured retry/error policies. Commits 37176cb, 7294115. |
| 2026-08-07 — 3.3.6 | Improved legacy iterator ordering and ensured worker cancellation/cleanup when iteration finishes or the consumer exits early. Commits 34286d0, 76b3023. |
| 2026-08-04 — 3.3.5 | Added local interface binding and corrected proxy initialization to use a single proxy URL. Commit a3ca95e. |
| 2026-07-27 — 3.3.4 | Hardened output filename sanitization against path traversal and invalid filename components. Commit eec47e3. |
| 2026-07-26 — 3.3.3 | Baseline requested for this documentation update; licensing changed to AGPL-3.0-or-later. Commit 3ee47f7. |