Skip to main content

使用python将html转化成json并部署到服务器

Project description

PHJF Alpha v0.0.3

  • 简单易懂的爬虫
  • 使用python将html转化成json并部署到服务器

>>注意<<: 如果想使用本库的全部内容,请至少达到能知道大多数Python3基础知识并且知道BeautifulSoup的基本用法和爬虫的概念和Flask库的基本知识

立志于用最简单的工作


为什么选择PHJF?

  • 简单
  • 轻便
  • 超强的可塑性
  • 一键部署

实现的功能

  • 模块化
  • 算是debug
  • 保存为json并且格式化
  • 键部署到本地服务器

未实现的功能

  • 部署到真正意义上的互联网服务器
  • 爬虫未响应自动结束程序
  • 部署到服务器时进行格式化

迅速开始

使用pip下载 pip3 install PHJF

注意:所依赖的第三方库

  • BeautifulSoup
  • Flask

库地址:PHJF in Github


快速入门

最简单的项目

创建main.py

注意:必须是main.py

get_page(url, encoding, 工作模式)

获取页面,拥有三种工作模式

  • url : 输入您所需要爬虫的网页
  • encoding : 编码格式,默认为 utf-8
  • 工作模式 : 见下文

工作模式的选择

  • let : 对接页面数据,解析时使用
  • save : 保存页面到当前目录

玩法实例

爬取baidu.com的数据并保存

from PHJF import *

def main():
  get_page("https://baidu.com", "", "save")

if __name__ == "__main__"
  main()

进阶玩法

from PHJf import *


def data(page_text, lists_info):
    soup = BeautifulSoup(page_text, 'lxml')
    # 在这里写入你需要的
    lists_info.append({})


def main():
    get_page("", "", "")
    run_compile_page("", "")
    run_server()


if __name__ == "__main__":
    main()

data()

不要改变data()里的内容

  • data 函数用来注入soup来编译html
  • lists_info.append({}) 用来放置输出,未来将会编译成中文json

比如

def data(page_text, lists_info):
    soup = BeautifulSoup(page_text, 'lxml')
    page_list = soup.select('.Revision_list > ul > li')
    for each in page_list:
        image = each.find("img")
        image_url = image['data-original']
        title = each.find("a", attrs={"class": "bt"})
        text = each.find("div", attrs={"class": "miaoshu"}).text
        data = each.find("span", attrs={"class": "time"}).text
        lists_info.append({"image": image_url, "title": title.text, "text": text, "data": data})

run_compile_page(工作模式, 文件名字)

工作模式

  • json : 保存页面到当前目录
  • data : 为运行本地服务器对接数据

文件名字

  • 为你的本地服务器设置目录名称与保存文件时的名称

run_server()

  • 启动本地服务器

常见问题

  • Q:爬虫爬不动了
    A:重新启动程序

*来自 邱璇洛 2022 ©*️

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

PHJF-0.0.4.tar.gz (3.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

PHJF-0.0.4-py3-none-any.whl (3.5 kB view details)

Uploaded Python 3

File details

Details for the file PHJF-0.0.4.tar.gz.

File metadata

  • Download URL: PHJF-0.0.4.tar.gz
  • Upload date:
  • Size: 3.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/4.0.1 CPython/3.8.5

File hashes

Hashes for PHJF-0.0.4.tar.gz
Algorithm Hash digest
SHA256 41900d5408cad70569c065e58833a4d3bc5ee005e50202c8b8356dac03b84996
MD5 b957ee32ef9014178d3d001c29f01a96
BLAKE2b-256 808e8952a7e6b44539d425ea04b51581bd43ce508e99d644c7c7b9050a151ad3

See more details on using hashes here.

File details

Details for the file PHJF-0.0.4-py3-none-any.whl.

File metadata

  • Download URL: PHJF-0.0.4-py3-none-any.whl
  • Upload date:
  • Size: 3.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/4.0.1 CPython/3.8.5

File hashes

Hashes for PHJF-0.0.4-py3-none-any.whl
Algorithm Hash digest
SHA256 8fd48a6896fd5015d338ad6210d1fadb0b5193b61539545a2e27e9d3fcb4970d
MD5 d4d23e7ad450aae11334e52570a74445
BLAKE2b-256 97bc6000e975029bea63b92de2a19b029abd7eb728211209f1c19c20aee8488e

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page