Hướng Dẫn Thực Hành

Error budget cho CAPTCHA: theo dõi độ tin cậy bằng SLO

Một team QA tại công ty outsourcing ở TP.HCM giám sát pipeline giải CAPTCHA cho khách hàng. SLO đặt ra là 95% giải thành công, tuần trước chỉ đạt 94,2%. Có phải sự cố cần báo động? Câu trả lời không nằm ở một con số đơn lẻ, mà ở error budget — phần trăm thất bại còn được phép "tiêu" trước khi độ tin cậy tụt dưới SLO, và tốc độ đang tiêu nó nhanh hay chậm. Có error budget, bạn trả lời câu hỏi đó bằng dữ liệu, và biết chính xác khi nào cần dừng deploy hoặc điều tiết request.

SLO, error budget và burn rate: khái niệm cốt lõi

Khái niệm Định nghĩa Ví dụ
SLO Tỷ lệ thành công mục tiêu Giải thành công 95%
Error budget Tỷ lệ thất bại được phép 5% tổng số lần giải có thể thất bại
Burn rate Tốc độ tiêu thụ error budget 2× nghĩa là ngân sách cạn trong nửa cửa sổ đo
Window (cửa sổ đo) Khoảng thời gian tính toán Rolling 24 giờ hoặc 7 ngày

Đọc burn rate: khi nào cần hành động

Burn rate Ý nghĩa Hành động
< 1,0 Tiêu error budget chậm hơn dự kiến Không cần hành động
1,0 Đúng đà, sẽ cạn đúng cuối cửa sổ đo Theo dõi sát
2,0 Ngân sách cạn trong nửa thời gian cửa sổ Điều tra nguyên nhân, giảm tải
≥5,0 Tiêu ngân sách rất nhanh Tạm dừng các tác vụ giải không thiết yếu

Ví dụ pipeline theo dõi giá trên Shopee, Lazada:

  • SLO 95%, cửa sổ 24 giờ, 10.000 lượt giải/ngày → error budget 500 lần thất bại.
  • 500 lần thất bại rải đều 24 giờ → burn rate 1,0, chỉ cần theo dõi sát.
  • 500 lần thất bại dồn trong 12 giờ đầu → burn rate 2,0, tạm dừng deploy mới và đổi tham số retry tới khi hạ nhiệt.

Python: theo dõi error budget khi giải CAPTCHA

Class dưới đây gom event thành công/thất bại vào một cửa sổ trượt, tự tính burn rate và bắn callback khi trạng thái đổi (healthywarningcriticalexhausted) — gắn thẳng vào hàm gọi in.php/res.php.

import time
import threading
from dataclasses import dataclass, field
from collections import deque
from enum import Enum

API_KEY = "YOUR_API_KEY"

class BudgetStatus(Enum):
    HEALTHY = "healthy"          # Budget > 50% remaining
    WARNING = "warning"          # Budget 10-50% remaining
    CRITICAL = "critical"        # Budget < 10% remaining
    EXHAUSTED = "exhausted"      # Budget depleted

@dataclass
class SLOConfig:
    """Service Level Objective configuration."""
    target_success_rate: float = 0.95  # 95%
    window_seconds: int = 86400        # 24 hours
    warning_threshold: float = 0.50    # Alert at 50% budget
    critical_threshold: float = 0.10   # Alert at 10% budget

@dataclass
class ErrorBudgetEvent:
    timestamp: float
    success: bool

class ErrorBudgetTracker:
    """Tracks error budget consumption for CAPTCHA solving."""

    def __init__(self, config: SLOConfig = SLOConfig()):
        self.config = config
        self._events: deque[ErrorBudgetEvent] = deque()
        self._lock = threading.Lock()
        self._callbacks: dict[BudgetStatus, list[callable]] = {
            status: [] for status in BudgetStatus
        }
        self._last_status = BudgetStatus.HEALTHY

    def on_status_change(self, status: BudgetStatus, callback: callable):
        """Register a callback for status transitions."""
        self._callbacks[status].append(callback)

    def record(self, success: bool):
        """Record a solve attempt."""
        now = time.monotonic()
        event = ErrorBudgetEvent(timestamp=now, success=success)

        with self._lock:
            self._events.append(event)
            self._prune(now)
            new_status = self._compute_status()

            if new_status != self._last_status:
                self._last_status = new_status
                for cb in self._callbacks.get(new_status, []):
                    try:
                        cb(self.get_report())
                    except Exception as e:
                        print(f"[BUDGET] Callback error: {e}")

    def _prune(self, now: float):
        """Remove events outside the window."""
        cutoff = now - self.config.window_seconds
        while self._events and self._events[0].timestamp < cutoff:
            self._events.popleft()

    def _compute_status(self) -> BudgetStatus:
        remaining = self.remaining_fraction
        if remaining <= 0:
            return BudgetStatus.EXHAUSTED
        if remaining < self.config.critical_threshold:
            return BudgetStatus.CRITICAL
        if remaining < self.config.warning_threshold:
            return BudgetStatus.WARNING
        return BudgetStatus.HEALTHY

    @property
    def total_events(self) -> int:
        with self._lock:
            return len(self._events)

    @property
    def success_count(self) -> int:
        with self._lock:
            return sum(1 for e in self._events if e.success)

    @property
    def failure_count(self) -> int:
        with self._lock:
            return sum(1 for e in self._events if not e.success)

    @property
    def current_success_rate(self) -> float:
        total = self.total_events
        return self.success_count / total if total > 0 else 1.0

    @property
    def error_budget_total(self) -> float:
        """Total allowed failures in the window."""
        total = self.total_events
        if total == 0:
            return 0
        return total * (1 - self.config.target_success_rate)

    @property
    def error_budget_remaining(self) -> float:
        """Remaining failure allowance."""
        return max(0, self.error_budget_total - self.failure_count)

    @property
    def remaining_fraction(self) -> float:
        """Fraction of error budget remaining (0.0 to 1.0)."""
        budget = self.error_budget_total
        if budget <= 0:
            return 1.0 if self.failure_count == 0 else 0.0
        return max(0, self.error_budget_remaining / budget)

    @property
    def burn_rate(self) -> float:
        """How fast the budget is being consumed (1.0 = normal, 2.0 = 2× faster)."""
        total = self.total_events
        if total == 0:
            return 0.0
        expected_failures = total * (1 - self.config.target_success_rate)
        if expected_failures == 0:
            return 0.0
        return self.failure_count / expected_failures

    def get_report(self) -> dict:
        return {
            "status": self._last_status.value,
            "slo_target": self.config.target_success_rate,
            "current_rate": round(self.current_success_rate, 4),
            "total_events": self.total_events,
            "successes": self.success_count,
            "failures": self.failure_count,
            "budget_total": round(self.error_budget_total, 1),
            "budget_remaining": round(self.error_budget_remaining, 1),
            "budget_remaining_pct": round(self.remaining_fraction * 100, 1),
            "burn_rate": round(self.burn_rate, 2),
        }

# --- Integration with solver ---

budget = ErrorBudgetTracker(SLOConfig(
    target_success_rate=0.95,
    window_seconds=3600,  # 1-hour window for demo
))

# Register alerts
budget.on_status_change(BudgetStatus.WARNING, lambda r:
    print(f"[ALERT] Budget warning: {r['budget_remaining_pct']}% remaining"))

budget.on_status_change(BudgetStatus.CRITICAL, lambda r:
    print(f"[ALERT] Budget critical: {r['budget_remaining_pct']}% remaining"))

budget.on_status_change(BudgetStatus.EXHAUSTED, lambda r:
    print(f"[ALERT] Budget EXHAUSTED — throttle new requests"))

def solve_with_budget(params: dict) -> str:
    """Solve CAPTCHA while tracking error budget."""
    import requests

    if budget._last_status == BudgetStatus.EXHAUSTED:
        raise RuntimeError("Error budget exhausted — solving paused")

    try:
        submit_params = {**params, "key": API_KEY, "json": 1}
        resp = requests.post(
            "https://ocr.captchaai.com/in.php", data=submit_params, timeout=30
        ).json()
        if resp.get("status") != 1:
            budget.record(False)
            raise RuntimeError(f"Submit: {resp.get('request')}")

        task_id = resp["request"]
        start = time.monotonic()
        while time.monotonic() - start < 180:
            time.sleep(5)
            poll = requests.get("https://ocr.captchaai.com/res.php", params={
                "key": API_KEY, "action": "get", "id": task_id, "json": 1,
            }, timeout=15).json()

            if poll.get("request") == "CAPCHA_NOT_READY":
                continue
            if poll.get("status") == 1:
                budget.record(True)
                return poll["request"]

            budget.record(False)
            raise RuntimeError(f"Solve: {poll.get('request')}")

        budget.record(False)
        raise RuntimeError("Timeout")

    except Exception:
        budget.record(False)
        raise

# Usage
for i in range(100):
    try:
        token = solve_with_budget({
            "method": "turnstile",
            "sitekey": "0x4XXXXXXXXXXXXXXXXX",
            "pageurl": "https://example.com",
        })
    except RuntimeError as e:
        if "exhausted" in str(e):
            print(f"Stopped at iteration {i}")
            break

print(budget.get_report())

solve_with_budget() chặn request mới ngay khi trạng thái là EXHAUSTED, thay vì âm thầm gửi tiếp và để lỗi chất chồng.

JavaScript: theo dõi error budget khi giải CAPTCHA

Cùng logic, viết cho worker Node.js hoặc script chạy trong trình duyệt — dùng private field (#) để tránh lộ state nội bộ.

class ErrorBudgetTracker {
  #events = [];
  #config;
  #callbacks = {};

  constructor(config = {}) {
    this.#config = {
      targetRate: config.targetRate || 0.95,
      windowMs: config.windowMs || 3600_000,
      warningThreshold: config.warningThreshold || 0.5,
      criticalThreshold: config.criticalThreshold || 0.1,
    };
    this.lastStatus = "healthy";
  }

  on(status, callback) {
    this.#callbacks[status] = this.#callbacks[status] || [];
    this.#callbacks[status].push(callback);
  }

  record(success) {
    const now = Date.now();
    this.#events.push({ time: now, success });
    this.#prune(now);

    const newStatus = this.#computeStatus();
    if (newStatus !== this.lastStatus) {
      this.lastStatus = newStatus;
      for (const cb of this.#callbacks[newStatus] || []) {
        cb(this.report());
      }
    }
  }

  #prune(now) {
    const cutoff = now - this.#config.windowMs;
    while (this.#events.length && this.#events[0].time < cutoff) {
      this.#events.shift();
    }
  }

  #computeStatus() {
    const frac = this.remainingFraction;
    if (frac <= 0) return "exhausted";
    if (frac < this.#config.criticalThreshold) return "critical";
    if (frac < this.#config.warningThreshold) return "warning";
    return "healthy";
  }

  get total() { return this.#events.length; }
  get successes() { return this.#events.filter((e) => e.success).length; }
  get failures() { return this.#events.filter((e) => !e.success).length; }
  get currentRate() { return this.total ? this.successes / this.total : 1; }

  get budgetTotal() {
    return this.total * (1 - this.#config.targetRate);
  }

  get budgetRemaining() {
    return Math.max(0, this.budgetTotal - this.failures);
  }

  get remainingFraction() {
    const bt = this.budgetTotal;
    if (bt <= 0) return this.failures === 0 ? 1 : 0;
    return Math.max(0, this.budgetRemaining / bt);
  }

  get burnRate() {
    const expected = this.total * (1 - this.#config.targetRate);
    return expected > 0 ? this.failures / expected : 0;
  }

  report() {
    return {
      status: this.lastStatus,
      currentRate: Math.round(this.currentRate * 10000) / 10000,
      total: this.total,
      failures: this.failures,
      budgetRemainingPct: Math.round(this.remainingFraction * 1000) / 10,
      burnRate: Math.round(this.burnRate * 100) / 100,
    };
  }
}

// Usage
const budget = new ErrorBudgetTracker({ targetRate: 0.95, windowMs: 3600_000 });

budget.on("warning", (r) => console.log(`[WARN] ${r.budgetRemainingPct}% budget left`));
budget.on("exhausted", (r) => console.log("[ALERT] Budget exhausted!"));

// Record results from your solver
budget.record(true);   // success
budget.record(false);  // failure
console.log(budget.report());

Cả hai bản expose burnRate/burn_rate — con số nên đưa vào dashboard hoặc alert, không chỉ riêng tỷ lệ thành công.

Khắc phục sự cố thường gặp

Vấn đề Nguyên nhân Cách xử lý
Error budget cạn quá nhanh SLO quá chặt so với thực tế Đặt SLO dựa trên dữ liệu lịch sử
Error budget gần như không tiêu SLO quá dễ dãi Siết SLO để buộc cải thiện độ tin cậy
Trạng thái nhảy qua lại liên tục Cửa sổ đo quá ngắn Dùng cửa sổ dài hơn (24h thay vì 1h)
Burn rate lệch ở lượng request thấp Quá ít sự kiện, phép tính bị nhiễu Yêu cầu số sự kiện tối thiểu trước khi tính
Bộ nhớ tracker tăng dần Event cũ không bị dọn Kiểm tra _prune/#prune chạy ở mỗi record()

Câu hỏi thường gặp

Burn rate bao nhiêu thì nên tạm dừng giải CAPTCHA tự động?

Từ 2,0 trở lên: ngân sách cạn nhanh gấp đôi dự kiến. Ở mức ≥5,0, tạm dừng tác vụ giải không thiết yếu và giữ thread cho các luồng ưu tiên cao.

Error budget khác gì so với SLA uptime thông thường?

SLA uptime đo hệ thống có "sống" hay không; error budget đo chất lượng từng lần giải trong một cửa sổ trượt. Service có thể uptime 100% nhưng vẫn đốt hết error budget nếu tỷ lệ giải tụt dưới SLO — hai chỉ số bổ trợ nhau.

SLO thực tế cho từng loại CAPTCHA nên đặt bao nhiêu?

Tùy loại: reCAPTCHA v2 thường đạt 90–95%, Turnstile cao hơn, CAPTCHA hình ảnh dao động hơn. Đo tỷ lệ thành công thực tế trước, rồi đặt SLO thấp hơn baseline đó 2–3%.

Có cần error budget riêng cho từng loại CAPTCHA không?

Có. reCAPTCHA chịu SLO 93%, Turnstile chịu 97%. Gộp chung một ngân sách sẽ che mất vấn đề riêng của từng loại.

Nâng gói CaptchaAI có giúp burn rate hạ nhiệt không?

Không trực tiếp. CaptchaAI tính phí theo thread (BASIC $15/tháng, 5 thread; STANDARD $30/tháng, 15 thread) — nâng gói tăng thông lượng song song, không tự sửa tỷ lệ thất bại. Burn rate cao do lỗi tăng thì kiểm tra tham số gửi lên, timeout và loại CAPTCHA trước.

Bài viết liên quan

Bước tiếp theo

Đo độ tin cậy pipeline giải CAPTCHA bằng số liệu thay vì cảm tính — lấy API key CaptchaAI và gắn error budget tracker vào luồng gọi in.php/res.php.

Hướng dẫn liên quan:

Os comentários estão desativados para este artigo.