Skip to main content

jupyter-data-fetch

从JupyterLab、Jupyter Notebook、Kaggle、Google Colab、VSCode网页版/code-server中抓取数据的示例

优点

  1. 通用性强,理论上全平台通用
  2. 无需中转服务器,能打开网页就能使用

安装

  1. uv pip install jupyter-data-fetch -U -i https://mirrors.aliyun.com/pypi/simple # Jupyter消息协议版
  2. uv pip install jupyter-data-fetch[playwright] -U -i https://mirrors.aliyun.com/pypi/simple # playwright网页自动化版

Jupyter消息协议版(通用)

  1. 根据Jupyter消息协议,模拟浏览器直接连接服务器进行代码的执行和获取,效率高
  2. 支持JupyterLab、Jupyter Notebook、Kaggle、Google Colab等
  3. 参考examples/message

Jupyter消息协议 + HTTP下载版(速度快)

  1. 在数据获取阶段,不通过网页展示提取数据,而是得到下载地址后HTTP下载
  2. 直接是二进制,不用base64/base85编码,文件小下载快。适合大文件。但每个网站都需要针对性调整
  3. 部分平台由于权限问题,服务器上临时文件可能需要手工删除
  4. 部分平台HTTP下载有流量限制。请换回通用版。如:SuperMind
  5. 部分平台使用blob传输文件。请换回通用版。如:Kaggle、Google Colab
  6. 参考examples/download

playwright网页自动化版(全能但低效)

  1. 网页自动化控制,通用性更高,额外支持VSCode网页版/code-server
  2. 暂时不支持的网站也可以定制开发
  3. 效率较低,因为多了网页渲染
  4. 参考examples/automation

使用方法

  1. examples下提供了示例
  2. 以joinquant为例,打开浏览器,登录研究环境,按F12或Ctrl+Shift+I打开开发者工具
  3. 搜索kernels,复制Cookie devtool.png
  4. 替换示例中COOKIE即可 ide.png
  5. 会自动从COOKIE提取用户ID,并更新SERVER_URL

最简示例

from jupyter_kernel_client import KernelClient

from jupyter_data_fetch.codec import TextCodec
from jupyter_data_fetch import extract_from_reply

# ... 省去部分代码。更多参考examples/message/joinquant.py

with KernelClient(server_url="https://www.joinquant.com/user/12345678901", token=None, headers=headers) as kernel:
    # 一定要保证缩进正确
    code = """
df = get_fundamentals(query(
        valuation, income
    ).filter(
        # 这里不能使用 in 操作, 要使用in_()函数
        valuation.code.in_(['000001.XSHE', '600000.XSHG'])
    ), date='2015-10-15')
"""
    reply = kernel.execute(TextCodec.generate_code(code, var_name='df'), store_history=False)
    print(reply)
    obj = TextCodec.decode(extract_from_reply(reply))
    print(obj)

常用API的封装

实际开发时并不会每次都手工构造code,会将函数封装。例如

# jupyter_data_fetch/wraps/jqdatasdk.py
from jupyter_data_fetch import LazyCodec, LazyDownloader


# ======== 使用coder解码数据 ============
# 调用示例 examples/message/jqdatasdk.py
def get_industry(security, date=None):
    code = f"""_ = get_industry({repr(security)}, {repr(date)})"""
    code = LazyCodec.generate_code(code, var_name='_')
    # print(code)
    reply = LazyCodec.execute(code, store_history=False)
    return LazyCodec.decode_from_reply(reply)


# ======== 使用downloader下载数据,遇到流量限制还是换回codec解码 ============
# 调用示例 examples/download/joinquant.py
def get_all_securities(types=[], date=None):
    code = f"""_ = get_all_securities({repr(types)}, {repr(date)})"""
    code = LazyDownloader.generate_code(code, var_name='_')
    # print(code)
    reply = LazyDownloader.execute(code, store_history=False)
    return LazyDownloader.reply_down_replace_load(reply, show_progress=True, dst=None, load=True)

参考jupyter_data_fetch/wraps/jqdatasdk.py

也可以封装更复杂的代码为简单函数,例如:jqresearch_query_client.py

自动登录并获取数据的完整示例

参考examples/experimental/cookie_playwright.py

核心代码

  1. TextCodec: 目前使用base85编解码器,使用字符串传输数据,压缩率高。如果字符串被截断,必须使用ImageCodec
  2. ImageCodec: 图片编解码器,使用图片传输数据,base64编码压缩率低
  3. generate_code生成可在Notebook单元格中运行的代码字符串,一定要指定需要获取的变量名var_name
  4. kernel.execute在服务段执行字符串代码,返回json对象
  5. extract_from_reply从json中提取数据
  6. decode字符串解码成对象

注意

  1. 由于各平台限制,generate_code生成的代码可能无法运行,可以复制到Notebook中测试
  2. python3.6问题太多,可以打开一个ipynb文件后,通过菜单更改内核为最新版
  3. 可以连接到已经打开的内核,只要提供kernel_id参数即可。参考ricequant.py示例
  4. Notebook中可以导入当前目录中py,但本项目直接使用当前目录是/,导致导入失败,通过指定kernel_id可解决

Metadata

Release files for jupyter-data-fetch 0.2.4

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for jupyter-data-fetch 0.2.4
File Size Uploaded
jupyter_data_fetch-0.2.4.tar.gz 13.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for jupyter-data-fetch 0.2.4
File Interpreter ABI Platform
jupyter_data_fetch-0.2.4-py3-none-any.whl Python 3 none any Details

Total release size: 30.4 kB

Release files / jupyter_data_fetch-0.2.4.tar.gz

Download URL jupyter_data_fetch-0.2.4.tar.gz
Size 13.2 kB
Tags Source
SHA-256 checksum
How to use checksums
059f53e28d32c8efd5a871ed6fb8a2623097695d75f39720749614a0955bef4f
BLAKE2b-256 checksum
How to use checksums
df9ae58824205444119ac966ca14e317379fcfe27f52e5e27abf95ca94497c72
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.9.25

Release files / jupyter_data_fetch-0.2.4-py3-none-any.whl

Download URL jupyter_data_fetch-0.2.4-py3-none-any.whl
Size 17.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0b199c04a12ba2c771fbb469bf8e571c637669d1c159d2878a32eb8be7fe02d6
BLAKE2b-256 checksum
How to use checksums
14a76eaf908e6031e3a3849ca4256a3fb985a690a4672d904a6335d06e86268e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.9.25

Release history Release notifications | RSS feed

0.3.0

2 release files

This release

0.2.4 This release

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page