Cơ sở dữ liệu về lương, bảng việc làm và cổng thông tin lao động của chính phủ bảo vệ dữ liệu lương thưởng bằng Cloudflare Turnstile và reCAPTCHA. CAPTCHA được kích hoạt khi truy vấn phạm vi lương theo vai trò, vị trí hoặc ngành — đặc biệt là trong quá trình thu thập dữ liệu hàng loạt trên nhiều chức danh công việc. Dưới đây là cách xử lý chúng.
Mẫu CAPTCHA trên Cổng thông tin lương
| Loại nguồn | CAPTCHA | Trình kích hoạt |
|---|---|---|
| Các trang so sánh lương | Cloudflare Turnstile | Truy vấn tìm kiếm lặp lại |
| Bộ lọc lương bảng công việc | reCAPTCHA v2 | Tra cứu lương nhiều lần |
| Thống kê lao động chính phủ | CAPTCHA hình ảnh | Yêu cầu tải xuống dữ liệu |
| Trang lương doanh nghiệp | Cloudflare Challenge | Lượt xem trang hàng loạt |
| Nền tảng khảo sát nhân sự | reCAPTCHA v3 | Gửi biểu mẫu |
Người thu thập dữ liệu tiền lương
import requests
import time
import re
from dataclasses import dataclass
@dataclass
class SalaryRecord:
title: str
location: str
min_salary: float
max_salary: float
median_salary: float
sample_size: int
source: str
class SalaryCollector:
def __init__(self, api_key):
self.api_key = api_key
self.session = requests.Session()
self.session.headers.update({
"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36"
})
def collect_salary_data(self, portal_url, job_title, location):
"""Search for salary data, solving CAPTCHAs as needed."""
response = self.session.get(portal_url, params={
"title": job_title,
"location": location
})
if self._is_turnstile_challenge(response):
response = self._solve_turnstile_and_retry(response, portal_url)
return self._parse_salary_data(response.text, portal_url)
def collect_bulk(self, portal_url, job_titles, locations):
"""Collect salary data for multiple job title + location combos."""
results = []
for title in job_titles:
for location in locations:
try:
data = self.collect_salary_data(
portal_url, title, location
)
results.extend(data)
# Respectful delay between requests
time.sleep(2)
except Exception as e:
print(f"Failed for {title} in {location}: {e}")
return results
def _is_turnstile_challenge(self, response):
return (
response.status_code == 403 or
"cf-turnstile" in response.text or
"challenges.cloudflare.com" in response.text
)
def _solve_turnstile_and_retry(self, response, url):
match = re.search(r'data-sitekey="(0x[^"]+)"', response.text)
if not match:
raise ValueError("Turnstile sitekey not found")
resp = requests.post("https://ocr.captchaai.com/in.php", data={
"key": self.api_key,
"method": "turnstile",
"sitekey": match.group(1),
"pageurl": url,
"json": 1
})
task_id = resp.json()["request"]
for _ in range(60):
time.sleep(3)
result = requests.get("https://ocr.captchaai.com/res.php", params={
"key": self.api_key,
"action": "get",
"id": task_id,
"json": 1
})
data = result.json()
if data["status"] == 1:
return self.session.post(url, data={
"cf-turnstile-response": data["request"]
})
raise TimeoutError("Turnstile solve timed out")
def _parse_salary_data(self, html, source):
from bs4 import BeautifulSoup
soup = BeautifulSoup(html, "html.parser")
records = []
def text_or_empty(node):
return node.text.strip() if node and node.text else ""
for row in soup.select(".salary-row, .compensation-entry, tr[data-salary]"):
try:
records.append(SalaryRecord(
title=text_or_empty(row.select_one(".job-title, .title")),
location=text_or_empty(row.select_one(".location")),
min_salary=self._parse_amount(
text_or_empty(row.select_one(".min-salary, .low"))
),
max_salary=self._parse_amount(
text_or_empty(row.select_one(".max-salary, .high"))
),
median_salary=self._parse_amount(
text_or_empty(row.select_one(".median, .mid"))
),
sample_size=int(
text_or_empty(row.select_one(".count, .sample")).replace(",", "") or 0
),
source=source
))
except (AttributeError, ValueError):
continue
return records
def _parse_amount(self, text):
if not text:
return 0.0
cleaned = re.sub(r'[^\d.]', '', text)
return float(cleaned) if cleaned else 0.0
# Usage
collector = SalaryCollector("YOUR_API_KEY")
data = collector.collect_bulk(
"https://salary.example.com/search",
job_titles=["Software Engineer", "Data Analyst", "Product Manager"],
locations=["San Francisco", "New York", "Austin"]
)
for record in data:
print(f"{record.title} in {record.location}: "
f"${record.min_salary:,.0f}–${record.max_salary:,.0f} "
f"(median: ${record.median_salary:,.0f})")
Tổng hợp nhiều nguồn (JavaScript)
class SalaryAggregator {
constructor(apiKey) {
this.apiKey = apiKey;
this.sources = [];
}
addSource(name, searchUrl) {
this.sources.push({ name, searchUrl });
}
async collectForRole(jobTitle, location) {
const results = [];
for (const source of this.sources) {
try {
const data = await this.querySource(source, jobTitle, location);
results.push({ source: source.name, ...data });
} catch (error) {
results.push({ source: source.name, error: error.message });
}
}
return this.aggregateResults(results, jobTitle, location);
}
async querySource(source, jobTitle, location) {
const url = `${source.searchUrl}?title=${encodeURIComponent(jobTitle)}&location=${encodeURIComponent(location)}`;
const response = await fetch(url);
const html = await response.text();
if (html.includes('cf-turnstile') || response.status === 403) {
return this.solveAndRetry(source.searchUrl, html, jobTitle, location);
}
return this.parseSalaryData(html);
}
async solveAndRetry(baseUrl, html, jobTitle, location) {
const match = html.match(/data-sitekey="(0x[^"]+)"/);
if (!match) throw new Error('Turnstile sitekey not found');
const submitResp = await fetch('https://ocr.captchaai.com/in.php', {
method: 'POST',
body: new URLSearchParams({
key: this.apiKey,
method: 'turnstile',
sitekey: match[1],
pageurl: baseUrl,
json: '1'
})
});
const { request: taskId } = await submitResp.json();
for (let i = 0; i < 60; i++) {
await new Promise(r => setTimeout(r, 3000));
const result = await fetch(
`https://ocr.captchaai.com/res.php?key=${this.apiKey}&action=get&id=${taskId}&json=1`
);
const data = await result.json();
if (data.status === 1) {
const response = await fetch(baseUrl, {
method: 'POST',
body: new URLSearchParams({
'cf-turnstile-response': data.request,
title: jobTitle,
location: location
})
});
return this.parseSalaryData(await response.text());
}
}
throw new Error('Turnstile solve timed out');
}
aggregateResults(results, jobTitle, location) {
const valid = results.filter(r => !r.error && r.median);
if (valid.length === 0) return null;
const medians = valid.map(r => r.median);
return {
jobTitle,
location,
avgMedian: medians.reduce((a, b) => a + b, 0) / medians.length,
sources: valid.length,
range: { min: Math.min(...medians), max: Math.max(...medians) }
};
}
}
// Usage
const aggregator = new SalaryAggregator('YOUR_API_KEY');
aggregator.addSource('SalaryDB', 'https://salarydb.example.com/search');
aggregator.addSource('PayScale', 'https://payscale.example.com/lookup');
const result = await aggregator.collectForRole('Software Engineer', 'San Francisco');
console.log(`Median salary: $${result.avgMedian.toLocaleString()} (${result.sources} sources)`);
Chiến lược thu thập dữ liệu
| Cách tiếp cận | Khối lượng mỗi ngày | tần số CAPTCHA | Tốt nhất cho |
|---|---|---|---|
| Tuần tự có độ trễ | 100–500 truy vấn | Thấp | Khảo sát nhỏ |
| Xoay vòng proxy | 500–2.000 truy vấn | Trung bình | Phân tích khu vực |
| Nhiều phiên song song | 2.000–10.000 truy vấn | Cao | Bộ dữ liệu toàn diện |
Khắc phục sự cố
| Vấn đề | Nguyên nhân | Cách xử lý |
|---|---|---|
| Cloudflare Turnstile trên mọi tìm kiếm | Phiên đã hết hạn | Duy trì cookie |
| Dữ liệu lương hiển thị "Yêu cầu đăng nhập" | Cổng thông tin yêu cầu xác thực | Xác thực trước khi tìm kiếm |
| Kết quả trống sau khi giải CAPTCHA | Thiếu tham số POST | Bao gồm tất cả các trường biểu mẫu ẩn |
| Dữ liệu không nhất quán giữa các lần chạy | Cổng thông tin hiển thị các phạm vi khác nhau | Sử dụng các tham số truy vấn nhất quán |
Câu hỏi thường gặp
Tôi có thể thực hiện bao nhiêu truy vấn về lương mỗi ngày?
Nó phụ thuộc vào giới hạn tốc độ của cổng chứ không phải CaptchaAI. CaptchaAI giải quyết Turnstile với tỷ lệ thành công 100%. Yêu cầu khoảng cách cách nhau 2–5 giây và xoay proxy để thu thập số lượng lớn.
Tôi có nên sử dụng proxy để thu thập dữ liệu tiền lương không?
Có, đặc biệt là để thu thập số lượng lớn hàng nghìn chức danh công việc. Proxy dân dụng giảm đáng kể tần suất CAPTCHA so với IP của trung tâm dữ liệu.
Tôi có thể thu thập dữ liệu tiền lương theo thời gian thực không?
Hầu hết các cổng lương đều cập nhật dữ liệu hàng tháng hoặc hàng quý nên việc thu thập dữ liệu theo thời gian thực là không cần thiết. Lên lịch chạy bộ sưu tập hàng tuần hoặc hàng tháng để có bộ dữ liệu toàn diện.
bài viết liên quan
- Thu thập dữ liệu nghiên cứu thị trường
- So sánh Geetest và Cloudflare Turnstile
- Cloudflare Turnstile 403 Sau khi sửa mã thông báo
Các bước tiếp theo
Thu thập dữ liệu bồi thường một cách đáng tin cậy —lấy khóa API CaptchaAI của bạnvà tự động xử lý CAPTCHA của cổng thông tin lương.