用Python爬取网页数据,从入门到实战。注意:爬虫请遵守网站robots.txt协议!
一、环境准备
# 安装依赖
pip install requests
pip install beautifulsoup4
pip install lxml
二、基础:requests获取网页
import requests
url = 'https://example.com'
headers = {
'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36'
}
response = requests.get(url, headers=headers)
print(response.status_code)
print(response.text)
三、解析HTML:BeautifulSoup
from bs4 import BeautifulSoup
soup = BeautifulSoup(response.text, 'lxml')
# 查找元素
title = soup.find('h1').text
links = soup.find_all('a')
# CSS选择器
items = soup.select('.item-class')
first = soup.select_one('#first-item')
四、实战:爬取新闻标题
import requests
from bs4 import BeautifulSoup
def scrape_news():
url = 'https://news.ycombinator.com'
response = requests.get(url)
soup = BeautifulSoup(response.text, 'html.parser')
titles = []
for item in soup.select('.titleline'):
title = item.find('a').text
link = item.find('a')['href']
titles.append({'title': title, 'link': link})
return titles
news = scrape_news()
for n in news[:5]:
print(f"标题: {n['title']}")
print(f"链接: {n['link']}")
print("---")
五、处理分页
import time
def scrape_all_pages():
base_url = 'https://example.com/page/{}'
all_data = []
for page in range(1, 6):
url = base_url.format(page)
response = requests.get(url)
soup = BeautifulSoup(response.text, 'lxml')
# 解析数据...
data = parse_page(soup)
all_data.extend(data)
time.sleep(1) # 礼貌延迟
return all_data
六、保存数据
# 保存为CSV
import csv
with open('data.csv', 'w', newline='', encoding='utf-8') as f:
writer = csv.DictWriter(f, fieldnames=['title', 'link'])
writer.writeheader()
writer.writerows(data)
# 保存为JSON
import json
with open('data.json', 'w', encoding='utf-8') as f:
json.dump(data, f, ensure_ascii=False, indent=2)
七、反爬处理
- 设置随机User-Agent
- 使用代理IP
- 控制请求频率
- 使用Selenium处理JS渲染
八、法律注意
重要提醒:爬虫可能违反网站服务条款。请遵守robots.txt,不要爬取敏感数据,不要影响网站正常运行。
总结
爬虫是数据分析的重要技能。掌握了基础后,可以学习Scrapy框架和异步爬虫。