jupyter-data-fetch
从JupyterLab、Jupyter Notebook、Kaggle、Google Colab、VSCode网页版/code-server中抓取数据的示例
优点
- 通用性强,理论上全平台通用
- 无需中转服务器,能打开网页就能使用
安装
uv pip install jupyter-data-fetch -U -i https://mirrors.aliyun.com/pypi/simple# Jupyter消息协议版uv pip install jupyter-data-fetch[playwright] -U -i https://mirrors.aliyun.com/pypi/simple# playwright网页自动化版
Jupyter消息协议版(通用)
- 根据Jupyter消息协议,模拟浏览器直接连接服务器进行代码的执行和获取,效率高
- 支持
JupyterLab、Jupyter Notebook、Kaggle、Google Colab等 - 参考examples/message
Jupyter消息协议 + HTTP下载版(速度快)
- 在数据获取阶段,不通过网页展示提取数据,而是得到下载地址后HTTP下载
- 直接是二进制,不用
base64/base85编码,文件小下载快。适合大文件。但每个网站都需要针对性调整 - 部分平台由于权限问题,服务器上临时文件可能需要手工删除
- 部分平台
HTTP下载有流量限制。请换回通用版。如:SuperMind - 部分平台使用
blob传输文件。请换回通用版。如:Kaggle、Google Colab - 参考examples/download
playwright网页自动化版(全能但低效)
- 网页自动化控制,通用性更高,额外支持
VSCode网页版/code-server - 暂时不支持的网站也可以定制开发
- 效率较低,因为多了网页渲染
- 参考examples/automation
使用方法
examples下提供了示例- 以
joinquant为例,打开浏览器,登录研究环境,按F12或Ctrl+Shift+I打开开发者工具 - 搜索
kernels,复制Cookie - 替换示例中
COOKIE即可 - 会自动从
COOKIE提取用户ID,并更新SERVER_URL
最简示例
from jupyter_kernel_client import KernelClient
from jupyter_data_fetch.codec import TextCodec
from jupyter_data_fetch import extract_from_reply
# ... 省去部分代码。更多参考examples/message/joinquant.py
with KernelClient(server_url="https://www.joinquant.com/user/12345678901", token=None, headers=headers) as kernel:
# 一定要保证缩进正确
code = """
df = get_fundamentals(query(
valuation, income
).filter(
# 这里不能使用 in 操作, 要使用in_()函数
valuation.code.in_(['000001.XSHE', '600000.XSHG'])
), date='2015-10-15')
"""
reply = kernel.execute(TextCodec.generate_code(code, var_name='df'), store_history=False)
print(reply)
obj = TextCodec.decode(extract_from_reply(reply))
print(obj)
常用API的封装
实际开发时并不会每次都手工构造code,会将函数封装。例如
# jupyter_data_fetch/wraps/jqdatasdk.py
from jupyter_data_fetch import LazyCodec, LazyDownloader
# ======== 使用coder解码数据 ============
# 调用示例 examples/message/jqdatasdk.py
def get_industry(security, date=None):
code = f"""_ = get_industry({repr(security)}, {repr(date)})"""
code = LazyCodec.generate_code(code, var_name='_')
# print(code)
reply = LazyCodec.execute(code, store_history=False)
return LazyCodec.decode_from_reply(reply)
# ======== 使用downloader下载数据,遇到流量限制还是换回codec解码 ============
# 调用示例 examples/download/joinquant.py
def get_all_securities(types=[], date=None):
code = f"""_ = get_all_securities({repr(types)}, {repr(date)})"""
code = LazyDownloader.generate_code(code, var_name='_')
# print(code)
reply = LazyDownloader.execute(code, store_history=False)
return LazyDownloader.reply_down_replace_load(reply, show_progress=True, dst=None, load=True)
参考jupyter_data_fetch/wraps/jqdatasdk.py
也可以封装更复杂的代码为简单函数,例如:jqresearch_query_client.py
自动登录并获取数据的完整示例
参考examples/experimental/cookie_playwright.py
核心代码
TextCodec: 目前使用base85编解码器,使用字符串传输数据,压缩率高。如果字符串被截断,必须使用ImageCodecImageCodec: 图片编解码器,使用图片传输数据,base64编码压缩率低generate_code生成可在Notebook单元格中运行的代码字符串,一定要指定需要获取的变量名var_namekernel.execute在服务段执行字符串代码,返回json对象extract_from_reply从json中提取数据decode字符串解码成对象
注意
- 由于各平台限制,
generate_code生成的代码可能无法运行,可以复制到Notebook中测试 python3.6问题太多,可以打开一个ipynb文件后,通过菜单更改内核为最新版- 可以连接到已经打开的内核,只要提供
kernel_id参数即可。参考ricequant.py示例 Notebook中可以导入当前目录中py,但本项目直接使用当前目录是/,导致导入失败,通过指定kernel_id可解决
Metadata
Release files for jupyter-data-fetch 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| jupyter_data_fetch-0.3.0.tar.gz | 13.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| jupyter_data_fetch-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 30.4 kB
Release files / jupyter_data_fetch-0.3.0.tar.gz
| Download URL | jupyter_data_fetch-0.3.0.tar.gz |
|---|---|
| Size | 13.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
c6e970478641bd2bf3cd9436bce0ea9ca9436acac747ddff267bf06d7dfb2c00
|
|
BLAKE2b-256 checksum How to use checksums |
08f9e12211cbd30a998f8fc9cae34d66df9b77d531048b45740430bd2f96eb56
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.25
|
Release files / jupyter_data_fetch-0.3.0-py3-none-any.whl
| Download URL | jupyter_data_fetch-0.3.0-py3-none-any.whl |
|---|---|
| Size | 17.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
41c5ebdef54c4a8a38beece9da4ceb929c0c5eaa82c8a58c2a3e132747105e30
|
|
BLAKE2b-256 checksum How to use checksums |
4394672de2ee475d6fc37734944f655d9f0b152be6ac5ccd8e57f6a796be84e9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.25
|