Fadey's Blog

arrow_back返回首页
Python爬虫

Python爬虫入门教程

用Python爬取网页数据,从入门到实战。注意:爬虫请遵守网站robots.txt协议!

一、环境准备

# 安装依赖
pip install requests
pip install beautifulsoup4
pip install lxml

二、基础:requests获取网页

import requests

url = 'https://example.com'
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36'
}

response = requests.get(url, headers=headers)
print(response.status_code)
print(response.text)

三、解析HTML:BeautifulSoup

from bs4 import BeautifulSoup

soup = BeautifulSoup(response.text, 'lxml')

# 查找元素
title = soup.find('h1').text
links = soup.find_all('a')

# CSS选择器
items = soup.select('.item-class')
first = soup.select_one('#first-item')

四、实战:爬取新闻标题

import requests
from bs4 import BeautifulSoup

def scrape_news():
    url = 'https://news.ycombinator.com'
    response = requests.get(url)
    soup = BeautifulSoup(response.text, 'html.parser')
    
    titles = []
    for item in soup.select('.titleline'):
        title = item.find('a').text
        link = item.find('a')['href']
        titles.append({'title': title, 'link': link})
    
    return titles

news = scrape_news()
for n in news[:5]:
    print(f"标题: {n['title']}")
    print(f"链接: {n['link']}")
    print("---")

五、处理分页

import time

def scrape_all_pages():
    base_url = 'https://example.com/page/{}'
    all_data = []
    
    for page in range(1, 6):
        url = base_url.format(page)
        response = requests.get(url)
        soup = BeautifulSoup(response.text, 'lxml')
        
        # 解析数据...
        data = parse_page(soup)
        all_data.extend(data)
        
        time.sleep(1)  # 礼貌延迟
    
    return all_data

六、保存数据

# 保存为CSV
import csv

with open('data.csv', 'w', newline='', encoding='utf-8') as f:
    writer = csv.DictWriter(f, fieldnames=['title', 'link'])
    writer.writeheader()
    writer.writerows(data)

# 保存为JSON
import json

with open('data.json', 'w', encoding='utf-8') as f:
    json.dump(data, f, ensure_ascii=False, indent=2)

七、反爬处理

  • 设置随机User-Agent
  • 使用代理IP
  • 控制请求频率
  • 使用Selenium处理JS渲染

八、法律注意

重要提醒:爬虫可能违反网站服务条款。请遵守robots.txt,不要爬取敏感数据,不要影响网站正常运行。

总结

爬虫是数据分析的重要技能。掌握了基础后,可以学习Scrapy框架和异步爬虫。