当前位置：首页 > news >正文

爬虫实战--人民网

news 2026/3/31 19:51:58

文章目录

前言
发现宝藏

前言

为了巩固所学的知识，作者尝试着开始发布一些学习笔记类的博客，方便日后回顾。当然，如果能帮到一些萌新进行新技术的学习那也是极好的。作者菜菜一枚，文章中如果有记录错误，欢迎读者朋友们批评指正。
（博客的参考源码可以在我主页的资源里找到，如果在学习的过程中有什么疑问欢迎大家在评论区向我提出）

发现宝藏

前些天发现了一个巨牛的人工智能学习网站，通俗易懂，风趣幽默，忍不住分享一下给大家。【宝藏入口】。

http://jhsjk.people.cn/testnew/result

import os
import re
from datetime import datetime
import requests
import json
from bs4 import BeautifulSoup
from pymongo import MongoClient
from tqdm import tqdmclass ArticleCrawler:def __init__(self, catalogues_url, card_root_url, output_dir, db_name='ren-ming-wang'):self.catalogues_url = catalogues_urlself.card_root_url = card_root_urlself.output_dir = output_dirself.client = MongoClient('mongodb://localhost:27017/')self.db = self.client[db_name]self.catalogues = self.db['catalogues']self.cards = self.db['cards']self.headers = {'Referer': 'https://jhsjk.people.cn/result?','User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) ''Chrome/119.0.0.0 Safari/537.36','Cookie': '替换成你自己的',}# 发送带参数的get请求并获取页面内容def fetch_page(self, url, page):params = {'keywords': '','isFuzzy': '0','searchArea': '0','year': '0','form': '','type': '0','page': page,'origin': '全部','source': '2',}response = requests.get(url, params=params, headers=self.headers)soup = BeautifulSoup(response.text, 'html.parser')return soup# 解析请求版面def parse_catalogues(self, json_catalogues):card_list = json_catalogues['list']for list in card_list:a_tag = 'article/'+list['article_id']card_url = self.card_root_url + a_tagcard_title = list['title']updateTime = list['input_date']self.parse_cards(card_url, updateTime)date = datetime.now()catalogues_id = list['article_id']+'01'# 检查重复标题existing_docs = self.catalogues.find_one({'id': catalogues_id})if existing_docs is not None:print(f'版面id: {catalogues_id}【已经存在】')continuecard_data = {'id': catalogues_id,'title': card_title,'page': 1,'serial': 1,# 一个版面一个文章'dailyId': '','cardSize': 1,'subjectCode': '50','updateTime': updateTime,'institutionnCode': '10000','date': date,'snapshot': {}}self.catalogues.insert_one(card_data)print(f'版面id: {catalogues_id}【插入成功】')# 解析请求文章def parse_cards(self, url, updateTime):response = requests.get(url, headers=self.headers)soup = BeautifulSoup(response.text, "html.parser")try:title = soup.find("div", "d2txt clearfix").find('h1').textexcept:try:title = soup.find('h1').textexcept:print(f'【无法解析该文章标题】{url}')html_content = soup.find('div', 'd2txt_con clearfix')text = html_content.get_text()imgs = [img.get('src') or img.get('data-src') for img in html_content.find_all('img')]cleaned_content = self.clean_content(text)# 假设我们有一个正则表达式匹配对象matchmatch = re.search(r'\d+', url)# 获取匹配的字符串card_id = match.group()date = datetime.now()if len(imgs) != 0:# 下载图片self.download_images(imgs, card_id)# 创建文档document = {'id': card_id,'serial': 1,'page': 1,'url' : url,'type': 'ren-ming-wang','catalogueId': card_id + '01','subjectCode': '50','institutionCode': '10000','updateTime': updateTime,'flag': 'true','date': date,'title': title,'illustrations': imgs,'html_content': str(html_content),'content': cleaned_content}# 检查重复标题existing_docs = self.cards.find_one({'id': card_id})if existing_docs is None:# 插入文档self.cards.insert_one(document)print(f"文章id：{card_id}【插入成功】")else:print(f"文章id：{card_id}【已经存在】")# 文章数据清洗def clean_content(self, content):if content is not None:content = re.sub(r'\r', r'\n', content)content = re.sub(r'\n{2,}', '', content)# content = re.sub(r'\n', '', content)content = re.sub(r' {6,}', '', content)content = re.sub(r' {3,}\n', '', content)content = content.replace('<P>', '').replace('<\P>', '').replace('&nbsp;', ' ')return content# 下载图片def download_images(self, img_urls, card_id):# 根据card_id创建一个新的子目录images_dir = os.path.join(self.output_dir, card_id)if not os.path.exists(images_dir):os.makedirs(images_dir)downloaded_images = []for img_url in img_urls:try:response = requests.get(img_url, stream=True)if response.status_code == 200:# 从URL中提取图片文件名image_name = os.path.join(images_dir, img_url.split('/')[-1])# 确保文件名不重复if os.path.exists(image_name):continuewith open(image_name, 'wb') as f:f.write(response.content)downloaded_images.append(image_name)print(f"Image downloaded: {img_url}")except Exception as e:print(f"Failed to download image {img_url}. Error: {e}")return downloaded_images# 如果文件夹存在则跳过else:print(f'文章id为{card_id}的图片文件夹已经存在')# 查找共有多少页def find_page_all(self, soup):# 查找<em>标签em_tag = soup.find('em', onclick=True)# 从onclick属性中提取页码if em_tag and 'onclick' in em_tag.attrs:onclick_value = em_tag['onclick']page_number = int(onclick_value.split('(')[1].split(')')[0])return page_numberelse:print('找不到总共有多少页数据')# 关闭与MongoDB的连接def close_connection(self):self.client.close()# 执行爬虫，循环获取多页版面及文章并存储def run(self):soup_catalogue = self.fetch_page(self.catalogues_url, 1)page_all = self.find_page_all(soup_catalogue)if page_all:for index in tqdm(range(1, page_all), desc='Page'):# for index in tqdm(range(1, 50), desc='Page'):soup_catalogues = self.fetch_page(self.catalogues_url, index).text# 解析JSON数据soup_catalogues_json = json.loads(soup_catalogues)self.parse_catalogues(soup_catalogues_json)print(f'======================================Finished page {index}======================================')self.close_connection()if __name__ == "__main__":crawler = ArticleCrawler(catalogues_url='http://jhsjk.people.cn/testnew/result',card_root_url='http://jhsjk.people.cn/',output_dir='D:\\ren-ming-wang\\img')crawler.run()  # 运行爬虫，搜索所有内容crawler.close_connection()  # 关闭数据库连接

在这里插入图片描述

爬虫实战--人民网

文章目录前言发现宝藏前言为了巩固所学的知识，作者尝试着开始发布一些学习笔记类的博客，方便日后回顾。当然，如果能帮到一些萌新进行新技术的学习那也是极好的。作者菜菜一枚，文章中如果有记录错误，欢迎读者朋友们…...

编程日记 2024/2/6 19:59:42

LGT8F328 UNO R3编译上传示例代码这是一段示例代码，将示例代码编译打包上传到LGT8F328 UNO R3开发板。 #include <Servo.h> Servo myservo; int pos 0; void setup() {// put your setup code here, to run once:Serial.begin(9600);Serial.println(&qu…...

编程日记 2024/2/6 19:58:41

Python进阶----在线翻译器（Python3的百度翻译爬虫）

目录一、此处需要安装第三方库requests: 二、抓包分析及编写Python代码 1、打开百度翻译的官网进行抓包分析。 2、编写请求模块 3、输出我们想要的消息三、所有代码如下： 一、此处需要安装第三方库requests: 在Pycharm平台终端或者命令提示符窗口中输入以下代…...

编程日记 2024/2/6 19:53:37

ArcGISPro中Python相关命令总结

主要总结conda方面的相关命令列出当前活动环境中的包 conda list 列出所有 conda 环境 conda env list 克隆环境克隆以默认的 arcgispro-py3 环境为模版的 my_env 新环境。 conda create --clone arcgispro-py3 --name my_env --pinned 激活环境 activate my_env p…...

编程日记 2024/2/6 19:51:35

2024年混合云：趋势和预测

混合云环境对于 DevOps 团队变得越来越重要，主要是因为它们能够弥合公共云资源的快速部署与私有云基础设施的安全和控制之间的差距。这种环境的混合为 DevOps 团队提供了灵活性和可扩展性，这对于大型企业中的持续集成和持续部署 (CI/CD) 至关重要。在混…...

编程日记 2024/2/6 19:50:34

c++入门学习④——对象的初始化和清理

目录对象的初始化和清理： why? 如何进行初始化和清理呢？ 使用构造函数和析构函数编辑构造函数语法: 析构函数语法: 构造函数的分类： 两种分类方式： 三种调用方法： 括号法（默认构造函数调用&…...

编程日记 2024/2/6 19:48:32

Java-spring注解的作用

1.Qualifier：通常与Autowired搭配使用，通过指定具体的beanName来注入相应的bean 当容器中有多个类型相同的Bean时，可以使用Qualifier注解来指定需要注入的Bean。Qualifier注解可以用于字段、方法参数、构造函数参数等位置 Service public cl…...

编程日记 2024/2/6 19:47:30

Allegro如何把Symbols，shapes，vias，Clines，Cline segs等多种元素一起移动

Allegro如何把Symbols，shapes，vias，Clines，Cline segs等多种元素一起移动在用Allegro进行PCB设计时，有时候需要同时移动某个区域的所有元素，如：Symbols，shapes，vias，Clines，Cline segs等元素。那么如何操作呢？首先就是把Symbols，shapes，vias，Clines，Cline …...

编程日记 2024/2/6 19:43:26

【力扣】罗马数字转整数，哈希集合+模拟

罗马数字转整数原题地址方法一：模拟罗马数字是字符串，其中每个字符都对应一个整数值，为了方便查找，可以预先把这种对应关系存储到哈希表中。遍历字符串，对于每个字符， 如果该字符不是最右边的字符&a…...

编程日记 2024/2/6 19:35:17

从长网址到短链接：探索网址缩短的神奇世界

欢迎来到我的博客，代码的世界里，每一行都是一个故事从长网址到短链接：探索网址缩短的神奇世界前言网址缩短的原理和历史网址缩短的应用场景网址缩短的安全风险网址缩短的未来趋势前言你是否曾经在浏览网页或社交媒体时遇到过一串看起来像…...

编程日记 2024/2/6 19:34:16

Micro micro controller一览

https://www.microchip.com.cn/， Microchip中文网站 https://www.microchip.com.cn/newcommunity/index.php?mSearch&adosearch&moduleDownload&keyworddsPIC33&p3 Microcontrollers and microProcessors dsPIC33 Digital Signal Controllers (D…...

编程日记 2024/2/6 19:31:12

一文简介Maven初级使用

一.概述 Maven是专门用于管理和构建Java项目的工具，它的主要功能有： 提供了一套标准化的项目结构提供了一套标准化的项目构建流程（编译，测试，打包，发布）提供了一套依赖管理机制一方面&…...

编程日记 2024/2/6 19:28:09

Django的配置文件setting.py

BASE_DIR 项目路径：默认是已经打开的主项目路径 BASE_DIR os.path.dirname(os.path.dirname(os.path.abspath(__file__))) SECRET_KEY 密钥 SECRET_KEY (dh&_fm2hfn9y)35!_6#$a7q%%^onoy#-a8x18r4(6*8f(aniDEBUG 帮助调试，默认…...

编程日记 2024/2/6 19:27:08

2024-02-06（Sqoop）

1.Sqoop Apache Sqoop是Hadoop生态体系和RDBMS（关系型数据库）体系之间传递数据的一种工具。 Sqoop工作机制是将导入或者导出命令翻译成MapReduce程序来实现。在翻译出的MapReduce中主要是对inputformat和outputformat进行定制。 Hadoop生态包括&#…...

编程日记 2024/2/6 19:22:02

C++ 11新特性之tuple

概述在C编程语言的发展历程中，C 11标准引入了许多开创性的新特性，极大地提升了开发效率与代码质量。其中，tuple（元组）作为一种强大的容器类型，为处理多个不同类型的值提供了便捷的手段。tuple是一种固定大…...

编程日记 2024/2/6 19:21:01

Spring Boot项目整合Seata AT模式

目录 1、添加依赖2.、配置Seata3、创建AT模式表4、使用Seata分布式事务 1、添加依赖 <dependency><groupId>io.seata</groupId><artifactId>seata-spring-boot-starter</artifactId></dependency>上述依赖适用于springboot项目如果你的项…...

编程日记 2024/2/6 19:17:58

作业2.5

第四章堆与拷贝构造函数一、程序阅读题 1、给出下面程序输出结果。 #include <iostream.h> class example {int a; public: example(int b5){ab;} void print(){aa1;cout <<a<<"";} void print()const {cout<<a<<endl;} …...

编程日记 2024/2/6 19:16:57

LeetCode、790. 多米诺和托米诺平铺【中等，二维DP，可转一维】

文章目录前言LeetCode、790. 多米诺和托米诺平铺【中等，二维DP，可转一维】题目与分类思路二维解法二维转一维资料获取前言博主介绍：✌目前全网粉丝2W，csdn博客专家、Java领域优质创作者，博客之星、阿里云平台优质…...

编程日记 2024/2/6 19:12:54

Python 的 sys 模块常用方法

sys.argv： 命令行参数 List，第一个元素是程序本身路径 sys.modules.keys()： 返回所有已经导入的模块列表 sys.exc_info() ：获取当前正在处理的异常类 exc_type、exc_value、exc_traceback 当前处理的异常详细信息 sys.exit(n)&…...

编程日记 2024/2/6 19:09:50

Kafka 使用手册

kafka3.0 文章目录 kafka3.01. 什么是kafka？2. kafka基础架构3. kafka集群搭建4. kafka命令行操作主题命令行【topic】生产者命令行【producer】消费者命令行【consumer】 5. kafka生产者生产者消息发送流程Producer 发送原理普通的异步发送带回调函数的异步发送同步…...

编程日记 2024/2/6 19:08:49