Skip to main content

mdsearch is a python package for paper search

Project description

代码调用

整体流程为通过search_info封装用户的查询内容,然后调用相关API返回检索结果。

  1. 初始化
from mdsearch import Searcher
S = Searcher(index_name='paperdb', doc_type='papers')
  1. 检索论文
paper, paper_id, paper_num = S.search_paper_by_name(search_info)
返回结果说明
paper_num   : A int number of paper.
paper_id    : A list of string, each string means paper id.
paper       : A list of dicts, each dict stores information of a paper.
  1. 检索单个论文视频中的相关内容(可通过前一步检索论文返回的paper_id或者paper,注意是单个)
video_pos = S.get_video_pos_by_paper_id(search_info, paper_id)
video_pos = S.get_video_pos_by_paper(search_info, paper)
  1. search_info 格式
# 综合检索
    search_info = {
        'query_type': 'integrated_search',
        'query': string,                                # 用户查询的内容
        'match': {
            'title': bool,                              # True/False表示是否检索这个字段的内容
            'abstract': bool,
            'paperContent': bool,
            'videoContent': bool,
        },
        'filter': {
            'yearfrom': 1000,                           # paper的年份限制
            'yearbefore': 3000,
        },
        # 'sort': 'relevance',
        'sort': 'year',                                 # 排序方式:year/cited/relevance
        'is_filter': False,                             # 是否先过滤后排序,建议True,提升检索效率
        'is_rescore': False,                            # 是否采用重排序,依据relevance排序时建议True,增强排序效果
        'is_cited': False                               # 是否使用引用量参与排序,由于未爬取引用量字段,只能为False
    }
# 高级检索
    search_info_2 = {
        'query_type': 'advanced_search',
        'match': {
            'title': string,                            # 用户查询内容
            'abstract': string,                         # 若不含有某项,设置成 None/False
            'paperContent': string,
            'videoContent': string,
        },
        'filter': {
            'yearfrom': 1000,
            'yearbefore': 3000,
        },
        'sort': 'year',
        'is_filter': False,                             # 高级检索无论True还是False都不会进行过滤
        'is_rescore': False,
        'is_cited': False
    }

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mdsearch-0.0.6.tar.gz (6.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

mdsearch-0.0.6-py3-none-any.whl (9.2 kB view details)

Uploaded Python 3

File details

Details for the file mdsearch-0.0.6.tar.gz.

File metadata

  • Download URL: mdsearch-0.0.6.tar.gz
  • Upload date:
  • Size: 6.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/3.3.0 pkginfo/1.6.1 requests/2.22.0 setuptools/51.1.2 requests-toolbelt/0.9.1 tqdm/4.51.0 CPython/3.6.4

File hashes

Hashes for mdsearch-0.0.6.tar.gz
Algorithm Hash digest
SHA256 43bd9b9ce4ecc5e9992184ec709add514cce2322cb5d15224fd91fafdd018068
MD5 7a1b227fc5980deaa93eb96f0ffdc522
BLAKE2b-256 55202d32fe5fff88a2dc423642c64611963c5f85721936185f004259a375a879

See more details on using hashes here.

File details

Details for the file mdsearch-0.0.6-py3-none-any.whl.

File metadata

  • Download URL: mdsearch-0.0.6-py3-none-any.whl
  • Upload date:
  • Size: 9.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/3.3.0 pkginfo/1.6.1 requests/2.22.0 setuptools/51.1.2 requests-toolbelt/0.9.1 tqdm/4.51.0 CPython/3.6.4

File hashes

Hashes for mdsearch-0.0.6-py3-none-any.whl
Algorithm Hash digest
SHA256 f7106bf4bfebc44b4d311361cc2403c2252eed6bd5ad0919905f7b8e7150d856
MD5 ac81f43e5ff122a79939430434e929e1
BLAKE2b-256 9a142daefd22c80319113c3df0b602da3ce71046f5147a4d30ae46dea4c85d6e

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page