python爬取贝壳中二手房的数据
·
前言:通过代码爬取贝壳中二手房的数据,以此给更多需要了解爬虫或者二手房信息的人提供便利。
第一部分:爬取地址
1.1贝壳首页地址
jiujiang.ke.com
第二部分:爬取数据
2.1输入要爬多少页
int(input('输入一共要多少页,输入整数:'))
2.2找到要请求的网址
wangzhi = f'https://jiujiang.ke.com/ershoufang/pg{y}/'
2.3定位多个房源信息位置
page.eles('xpath://div[@class="info clear"]', timeout=20)
2.4设置要爬的信息
for x in loc:
# 标题
biaoti = x.ele('xpath:.//div[@class="title"]/a').text
# 地址
address = x.ele('xpath:.//div[@class="positionInfo"]/a').text
# 概述
info = x.ele('xpath:.//div[@class="houseInfo"]').text
# 总价
price = x.ele('xpath:.//div[@class="totalPrice totalPrice2"]/span').text
# 详情页链接
lianjie = x.ele('xpath:.//div[@class="positionInfo"]/a').attr('href')
第三部分:保存信息
pandas.DataFrame(kong_list).to_csv('贝壳二手房(九江).csv', index=False)
全部的代码见文章末尾
更多信息请关注......,本篇文章的源码如下
from DrissionPage import ChromiumPage
import time
import pandas
kong_list = []
num = int(input('输入一共要多少页,输入整数:'))
for y in range(1, num+1):
page = ChromiumPage()
wangzhi = f'https://jiujiang.ke.com/ershoufang/pg{y}/'
page.get(wangzhi)
# 遇到需要验证的页码,就跳过进入下一页
loc = page.eles('xpath://div[@class="info clear"]', timeout=20) # 定位多个房源信息的位置
for x in loc:
# 标题
biaoti = x.ele('xpath:.//div[@class="title"]/a').text
# 地址
address = x.ele('xpath:.//div[@class="positionInfo"]/a').text
# 概述
info = x.ele('xpath:.//div[@class="houseInfo"]').text
# 总价
price = x.ele('xpath:.//div[@class="totalPrice totalPrice2"]/span').text
# 详情页链接
lianjie = x.ele('xpath:.//div[@class="positionInfo"]/a').attr('href')
print('\n', biaoti, address, info, price, lianjie)
kong_list.append({
'标题': biaoti,
'地址': address,
'概述': info,
'总价': price+'万',
'详情链接': lianjie,
})
time.sleep(3)
# 写入文件
pandas.DataFrame(kong_list).to_csv('贝壳二手房(九江).csv', index=False)
更多推荐



所有评论(0)