jupyter-data-fetch
从JupyterLab、Jupyter Notebook、VSCode网页版/code-server中抓取数据的示例
优点
- 与
ksrpc比,通用性更强,理论上全平台通用 - 不需中转服务器,网页能打开就能使用
安装
uv pip install jupyter-data-fetch -U# Jupyter消息协议版uv pip install jupyter-data-fetch[playwright] -U# playwright网页自动化版
Jupyter消息协议版(推荐)
- 根据Jupyter消息协议,模拟浏览器直接连接服务器进行代码的执行和获取,效率高
- 支持
JupyterLab、Jupyter Notebook - 参考examples/message
playwright网页自动化版
- 网页自动化控制,通用性更高,支持
JupyterLab、Jupyter Notebook - 支持
VSCode网页版/code-server - 暂时不支持的网站也可以定制开发
- 效率较低,因为多了网页渲染
- 参考examples/automation
使用方法
examples下提供了示例- 以
joinquant为例,打开浏览器,登录研究环境,按F12打开开发者工具 - 搜索
kernels,复制Cookie - 替换示例中
COOKIE即可 - 会自动从
COOKIE提取用户ID,并更新SERVER_URL
最简示例
from jupyter_kernel_client import KernelClient
from jupyter_data_fetch.codec import TextCodec, extract_from_reply
# ... 省去部分代码。更多参考examples/message/joinquant.py
with KernelClient(server_url="https://www.joinquant.com/user/12345678901", token=None, headers=headers) as kernel:
# 一定要保证缩进正确
code = """
df = get_fundamentals(query(
valuation, income
).filter(
# 这里不能使用 in 操作, 要使用in_()函数
valuation.code.in_(['000001.XSHE', '600000.XSHG'])
), date='2015-10-15')
"""
reply = kernel.execute(TextCodec.generate_code(code, var_name='df'), store_history=False)
print(reply)
obj = TextCodec.decode(extract_from_reply(reply))
print(obj)
自动登录并获取数据的完整示例
参考examples/experimental/cookie_playwright.py
核心代码
TextCodec: 目前使用base85编解码器,使用字符串传输数据,压缩率高。如果字符串被截断,必须使用ImageCodecImageCodec: 图片编解码器,使用图片传输数据,base64编码压缩率低generate_code生成可在Notebook单元格中运行的代码字符串,一定要指定需要获取的变量名var_namekernel.execute在服务段执行字符串代码,返回json对象extract_from_reply从json中提取数据decode字符串解码成对象
注意
- 由于各平台限制,
generate_code生成的代码可能无法运行,可以复制到Notebook中测试 python3.6问题太多,可以打开一个ipynb文件后,通过菜单更改内核为最新版- 可以连接到已经打开的内核,只要提供
kernel_id参数即可。参考ricequant.py示例 Notebook中可以导入当前目录中py,但本项目直接使用当前目录是/,导致导入失败,通过指定kernel_id可解决
Metadata
Release files for jupyter-data-fetch 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| jupyter_data_fetch-0.2.0.tar.gz | 10.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| jupyter_data_fetch-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 22.9 kB
Release files / jupyter_data_fetch-0.2.0.tar.gz
| Download URL | jupyter_data_fetch-0.2.0.tar.gz |
|---|---|
| Size | 10.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
2d615ff53c1a5aacc6009ac1cd9ff2e76ebd5db5fac437ce48aa257af682538e
|
|
BLAKE2b-256 checksum How to use checksums |
d5adbda2611c705b65dde9f9aa69f941a0404860651ed1900ac7e704c5f151d7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.25
|
Release files / jupyter_data_fetch-0.2.0-py3-none-any.whl
| Download URL | jupyter_data_fetch-0.2.0-py3-none-any.whl |
|---|---|
| Size | 12.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
6e34b51a96187fb660cfc09d105cc9ec9a94abc193df391b0513cdb00b655158
|
|
BLAKE2b-256 checksum How to use checksums |
4e4640ff9f16d6867b397c958242b8d31396c9e844602ae83af1d1f3d3a4d5b9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.25
|