Skip to main content

uiautomation_new - Refactored Python UIAutomation for Windows

uiautomation-new 后续不再更新,转到 winuia-auto uiautomation_new is a refactored Windows UI Automation package based on the classic uiautomation API. It keeps the familiar import style and most legacy function names while adding a cleaner internal architecture, modern input helpers, safer clicking, tab-window utilities, visual matching helpers, and PyPI-ready packaging metadata.

This package is designed for automating Windows desktop applications that expose Microsoft UIAutomation providers, including Win32, WinForms, WPF, Modern UI, Qt-based apps, browser windows, and Electron-based apps when accessibility support is enabled.

Installation

pip install uiautomation-new

The default installation includes comtypes, pywin32, Pillow, numpy, opencv-python, and pynput. Image matching and hotkey helpers work without installing feature extras. This project does not depend on pyautogui.

Development tools remain optional:

pip install "uiautomation-new[dev]"

Compatibility

The public compatibility layer is intentionally preserved:

import uiautomation as auto

window = auto.GetForegroundControl()
button = window.ButtonControl(Name="OK")
button.Click()

Most old names such as Click, SendKeys, ControlType, PatternId, PropertyId, WindowControl, ButtonControl, and EditControl remain available from the top-level package. New code can also use the added helper APIs described below.

What Changed in This Refactor

Modular Architecture

The old project was centered around one very large uiautomation.py file. This refactor starts splitting responsibilities into focused modules:

  • uiautomation/enums.py: UIAutomation and Win32 integer enums such as ControlType, PatternId, PropertyId, Keys, MouseEventFlag, and KeyboardEventFlag.
  • uiautomation/structures.py: Rect and ctypes input structures such as INPUT, MOUSEINPUT, KEYBDINPUT, and HARDWAREINPUT.
  • uiautomation/inputs/core.py: low-level SendInput, MouseInput, KeyboardInput, and virtual-key scan-code helpers.
  • uiautomation/inputs/keyboard.py: pynput hotkey helpers.
  • uiautomation/win32_api.py: pywin32-first, ctypes-fallback wrappers for selected Win32 APIs.
  • uiautomation/client.py: thread-safe singleton helpers and package DLL path handling.

The original uiautomation.py remains as a compatibility aggregation layer while larger areas such as controls, patterns, bitmap handling, and logging are migrated progressively.

Modernized Constants and Rect

  • Classic constant classes have been moved to enum.IntEnum while keeping int-compatible behavior.
  • Rect is now a dataclass-compatible helper in structures.py.
  • Empty rectangle detection now treats zero or negative width/height as empty.

Modern Input Simulation

  • Mouse clicks, mouse button down/up, mouse wheel, single-key input, and SendKeys now use SendInput internally.
  • Deprecated mouse_event and keybd_event names are still available as compatibility wrappers.
  • SendInput batches input events instead of sending them one by one.

Human-like Input Simulation

The standard input APIs can opt into human-like timing and movement with humanLike=True. This keeps the existing API surface while adding smoother mouse paths, randomized click timing, randomized typing intervals, and more natural wheel scrolling when explicitly requested.

import uiautomation as auto

# Move with a cubic Bezier path by default.
auto.MoveTo(500, 300, humanLike=True, duration=0.8)

# Click with human-like movement, pre-click pause, and mouse-down hold time.
auto.Click(500, 300, humanLike=True, duration=0.8)
control.Click(humanLike=True, duration=0.8, jitter=4)

# Type with variable intervals. duration controls how long the mouse takes to
# reach the control before typing. errorRate is optional and defaults to 0.
control.SendKeys(
    "hello from uiautomation",
    humanLike=True,
    duration=0.8,
    intervalRange=(0.05, 0.18),
    errorRate=0.01,
)

# Scroll with non-uniform wheel intervals and a controlled mouse-arrival time.
control.WheelDown(wheelTimes=5, humanLike=True, duration=0.8, intervalRange=(0.1, 0.3))

MoveTo() accepts curveFunc for advanced callers that need a custom movement curve. The function receives (startX, startY, targetX, targetY, t) and returns (x, y) for the current progress value:

def linear_curve(start_x, start_y, target_x, target_y, t):
    return (
        start_x + (target_x - start_x) * t,
        start_y + (target_y - start_y) * t,
    )

auto.MoveTo(500, 300, humanLike=True, duration=0.8, curveFunc=linear_curve)

Use duration=seconds when the mouse should take a fixed amount of time to reach the target. Use durationRange=(minSeconds, maxSeconds) when the arrival time should be randomized. duration takes priority when both are supplied.

SafeClick(humanLike=True) skips silent UIA invoke/post-message paths and uses the physical mouse click path. Use these options only for normal desktop UI automation where realistic input timing is desired; they are not intended for bypassing access controls or verification challenges.

DPI and Multi-Monitor Improvements

New helpers improve high-DPI and multi-monitor coordinate behavior:

import uiautomation as auto
auto.SetDpiAwareness()
rect = auto.GetVirtualScreenRect()
inside = auto.IsPointInVirtualScreen(x, y)

SetDpiAwareness() prefers PerMonitorV2 awareness and automatically falls back to older Windows DPI APIs.

Safer Clicks

Controls now support visibility and clickability checks:

control.IsVisibleOnScreen()
control.IsClickable()
control.SafeClick()
auto.SafeClick(x, y)

control.SafeClick() now prefers UIA InvokePattern for center-click style calls. This is more reliable for buttons whose visual center is covered by another layer or whose physical clickable point is inaccurate. If InvokePattern is unavailable or fails, it falls back to the previous mouse-click path with ScrollIntoView and clickability/occlusion checks.

auto.SafeClick(x, y) is a coordinate-based safe click: it gets the control at (x, y), tries silent UIA actions on the hit control and its ancestors, then tries a Win32 PostMessage click, and falls back to auto.Click(x, y) only when those silent attempts fail. PostMessage can only confirm that mouse messages were posted; some modern or hardware-accelerated apps may ignore them.

To force the physical mouse-click path:

control.SafeClick(preferInvoke=False)
auto.SafeClick(x, y, useInvoke=False, useLegacy=False, usePostMessage=False)

Controls can be filtered by relative position:

import uiautomation as auto
window = auto.ControlFromHandle(handle) # handle为窗口句柄
save = window.ButtonControl(Name="Save").RightOf(cancel_button)
save = window.ButtonControl(Name="Save").right_of(cancel_button)

Available helpers:

  • LeftOf() / left_of()
  • RightOf() / right_of()
  • Above() / above()
  • Below() / below()
  • AddCompare()
def relative_position():
    # 先定位一个已知的参考控件(作为“锚点”)
    reference = auto.ButtonControl(Name="智能分流")
    # 然后使用相对位置查找目标 A.RightOf(B) 表示“A 在 B 的右边”
    target = auto.ButtonControl(Name="全局代理").RightOf(reference)
    print(target)
    # 或者写成链式:
    target1 = reference.LeftOf(auto.TextControl(Name="全局代理"))  # 反过来也行?
    # # 注意:通常的写法是:目标控件的条件 .RightOf(参考控件)
    # 目标(智能分流) 在 参照物(全局代理)的左侧
    print(target1)
    #目标(快速连接) 在 参照物(智能分流)的上方
    # target2 = auto.TextControl(RegexName="快速连接").Above(reference)
    # print(target2)

Smart Waiting

Modern wait aliases are available:

control.WaitUntilExists(timeout=10)
control.WaitUntilDisappear(timeout=5)

auto.WaitUntilExists(control, timeout=10)
auto.WaitUntilDisappear(control, timeout=5)

UI Automation Event Listening

EventHub can listen to UI Automation events such as focus changes, window opened events, and property changes. Event delivery depends on COM message pumping; EventHub(..., autoPump=True) starts a background pump thread. Keep callbacks lightweight.

Typical use cases are places where event-driven waiting is cleaner than polling:

  • Wait for a login, save, or export dialog to open before finding its controls.
  • Wait for a button, menu item, or text box to become enabled before clicking or typing.
  • Record focus changes while debugging or building a simple script recorder.
import time
import uiautomation as auto
from uiautomation import uiautomation as auto_core

hub = auto.EventHub(auto_core._AutomationClient.instance().IUIAutomation)

def on_focus(sender):
    control = auto.Control.CreateControlFromElement(sender)
    print(control.ControlTypeName, control.Name)

with hub.listen_focus_changed(on_focus):
    time.sleep(10)

Property waits can avoid polling when the current value is not ready:

value = auto.WaitForProperty(
    button,
    auto.PropertyId.IsEnabledProperty,
    lambda current: bool(current),
    timeout=10,
    hub=hub,
)

Recursive FindAll and Tab Utilities

Windows with tab controls can be automated with FindAll() and tab helpers:

tabs = window.FindAll(auto.ControlType.TabItemControl, maxDepth=10)
for tab in tabs:
    url = auto.GetTabUrl(tab, window=window)
    print(tab.Name, url)

current_tab = window.GetCurrentTab()
window.SwitchToTabByName("Settings")
window.SwitchToTabByUrl("example.com")
url = window.GetCurrentUrl()
tab_urls = window.GetAllTabUrls()

Module-level helpers are also available:

tabs = auto.FindAll(window, auto.ControlType.TabItemControl)
tabs = auto.find_control_num(window, auto.ControlType.TabItemControl)
tab = auto.get_current_tab(window)
auto.switch_to_tab_by_name(window, "Settings")
auto.SwitchToTabByUrl(window, "example.com")
url = auto.get_current_url(window)
url = auto.GetTabUrl(tab, window=window)
tab_urls = auto.GetAllTabUrls(window)

FindAll() supports maxDepth to avoid walking extremely large UI trees indefinitely. Current-tab detection first uses SelectionItemPattern.IsSelected, then falls back to keyboard-focus properties. Tab switching prefers SelectionItemPattern.Select() and falls back to Click(). SwitchToTabByUrl() supports substring URL matching by default. GetTabUrl() and GetAllTabUrls() switch tabs to read the active address bar and then restore the originally selected tab when possible.

Chromium Browser Facade

Version 0.0.7 adds Browser, a high-level facade for a caller-provided Chromium window. It reuses the existing tab and URL helpers, locates localized address bars, and waits for a navigation change followed by stable UIA state.

browser = auto.Chrome()
browser.Navigate("https://example.com")
browser.WaitPageLoad(timeout=30)

def on_focus_changed(sender):
    print("focused:", sender.Name)

handler = auto.CreateFocusChangedEventHandler(on_focus_changed)
auto.AddFocusChangedEventHandler(handler)
try:
    # Browser/UI Automation work...
    pass
finally:
    auto.RemoveFocusChangedEventHandler(handler)

WaitPageLoad() waits for the URL/title/document UIA state to change and then remain stable for consecutive samples. It does not guarantee that every background network request has completed. For automatic cleanup, use the context-manager form:

with browser.ListenFocusChanged(on_focus_changed):
    # Focus events are delivered while this block is active.
    pass

Safe Top-Level Window Enumeration

Version 0.1.5 adds GetTopLevelWindowHandles() and GetTopLevelControls(). They enumerate top-level windows through Win32 EnumWindows before optionally converting each HWND to a UI Automation control. Prefer these helpers when enumerating desktop windows instead of walking GetRootControl().GetChildren(); individual UIA providers can still fail during ControlFromHandle().

handles = auto.GetTopLevelWindowHandles()
windows = auto.GetTopLevelControls()

Non-UIA Windows

NonUIAWindow is a separate HWND-based facade for windows that do not expose reliable UI Automation controls. Use WindowControl and other Control classes for UIA-capable applications; use NonUIAWindow for window-level automation, screenshots, template matching, OCR, and screen-coordinate input.

window = auto.NonUIAWindow(
    title_keyword="My App",
    class_name="MyWindowClass",
)

window.activate()
window.click(120, 48)  # relative to the window's top-left corner
window.send_keys("hello")
window.screenshot("window.png")

match = window.find_image("save.png", confidence=0.9)
window.click_image("save.png", timeout=5)

Pass hwnd when it is already known; it takes priority over keywords. Without an HWND, title_keyword and class_name each match by exact value or substring. A lookup that matches zero or multiple windows raises LookupError.

to_bitmap() tries PrintWindow first and falls back to a Win32/GDI screen capture when needed. Image matching is automatically limited to the current window rectangle. click_image() validates the latest window bounds before clicking the matched center.

close() sends WM_CLOSE by default, so an application can close its window while keeping background processes alive. Pass terminateProcess=True to force-stop the owning process; add killProcessTree=True to stop its child processes too. Forced termination does not save unsaved application data.

window.close()
window.close(terminateProcess=True, killProcessTree=True)

# Stop a known application root process and all of its descendants.
auto.TerminateProcessTree(root_process_id)

If an application destroys its HWND after hide() or set_topmost(), those methods return False instead of reporting a successful state change.

OCR is optional:

pip install "uiautomation-new[ocr]"
match = window.find_text("OK", minScore=0.8, offscreen=True)
window.click_text("OK", minScore=0.8, offscreen=True)

If a window later exposes usable UIA controls, conversion is explicit and optional:

uia_window = window.to_uia_control()
if uia_window:
    uia_window.ButtonControl(Name="OK").Click()
import uiautomation as auto

window = auto.WindowControl(
    ClassName="Chrome_WidgetWin_1",
    RegexName=".*Google Chrome",
)
browser = auto.Browser(window)

browser.navigate("https://example.com", timeout=30)
print(browser.get_current_url())
print(browser.get_current_tab().Name)

browser.select_tab("Example Domain")
browser.select_tab_by_url("example.com")
for tab in browser.get_all_tab_urls():
    print(tab["index"], tab["name"], tab["url"])

Browser currently targets Chromium windows such as Chrome and Edge. Pass the window explicitly to avoid attaching to an unintended browser instance. Page loading is inferred from URL, title, and document state; it is not a guarantee that a web application has completed every background request. Renderer content still requires the browser accessibility configuration described below.

P1 also provides factories that filter the matching Chromium window by process name, plus browser-standard history commands:

chrome = auto.Chrome(title="Example Domain")
edge = auto.Edge(title="Example Domain")

chrome.back()
chrome.forward()
chrome.reload()
document = chrome.get_current_document()

The factories raise LookupError when the requested executable is not running or no window title matches. get_current_document() prefers a visible DocumentControl; renderer content requires --force-renderer-accessibility. When no browser history entry exists, back() and forward() return False after the default one-second no_change_timeout rather than waiting for the full navigation timeout. reload() waits for stable UIA state even when the URL does not change; use wait=False when only issuing the shortcut matters.

The real browser integration test starts an isolated temporary profile and is disabled by default:

$env:UIAUTOMATION_RUN_BROWSER_INTEGRATION = "1"
$env:UIAUTOMATION_BROWSER = "chrome" # or "edge" or "firefox"
python tests\test_browser_integration.py

Firefox Profile and Page Extraction

Firefox can use the same facade after locating its top-level window and passing FIREFOX_PROFILE. Its address-bar profile supports the standard labels and the variable search-engine label used by Firefox.

firefox_window = auto.WindowControl(
    ClassName="MozillaWindowClass",
    SubName="Example Domain",
)
firefox = auto.Browser(firefox_window, profile=auto.FIREFOX_PROFILE)

Extract UIA content from the active document:

links = browser.get_all_links()
for link in links:
    print(link.Name)

text = browser.get_all_text()
print(text)

get_all_text() prefers TextPattern and falls back to visible TextControl names. Chromium may expose those controls through the window tree rather than as direct document children; the facade filters that fallback to controls with a DocumentControl ancestor. UIA text is not a DOM serialization and may omit content that the browser accessibility provider does not expose.

Browser Tabs and JavaScript Modes

Use browser-standard shortcuts to create or close tabs. new_tab(url) opens a URL in the newly active tab. close_tab() closes the active tab by default, closes one zero-based index, or closes every tab matching title and URL filters.

browser.new_tab()
browser.new_tab("https://example.com")
browser.close_tab()
browser.close_tab(2)
browser.close_tab(
    "Example Domain",
    url="https://example.com/orders/42",
    exact=True,
)

For close_tab(), an integer tab is always treated as the tab index and ignores url, so it closes only that tab. A non-empty string tab closes all title matches; add a non-empty url to require both values. Pass tab="" with a URL to close all URL matches. Matching tabs are closed in descending index order so index changes after each close do not affect later matches. Passing both tab="" and url="" raises ValueError.

For explicit page-side automation, the default JavaScript mode enters a javascript: URL through the address bar:

browser.execute_javascript("document.title = 'Ready'")

Use method="console" when the script must be issued through DevTools. It focuses the Console prompt and types the script directly, avoiding DevTools paste protection. If this call opens DevTools, it closes it after execution by default; set close_console=False to leave it open. A Console that was already open before the call is never toggled or closed.

browser.execute_javascript(
    "document.title = 'Ready'",
    method="console",
    close_console=True,
)

Set paste=True to use the text clipboard instead. The facade restores the prior text clipboard value before returning. Some DevTools sessions block pasting until the user enters allow pasting; pass allow_pasting=True to attempt that input first:

browser.execute_javascript(
    "document.title = 'Ready'",
    method="console",
    paste=True,
    allow_pasting=True,
)

Chrome may intentionally reject an automated allow pasting entry as part of its self-XSS protection. In that case, type the phrase once in the DevTools console manually, then call the paste=True mode. Neither mode retrieves a JavaScript return value or bypasses browser security policy. The caller is responsible for the script's effects.

pynput Hotkey Helper

auto.Press_Hotkey("ctrl", "a")
auto.Press_Hotkey("backspace")
auto.Press_Hotkey("ctrl", "shift", "a")
auto.Press_Hotkey("shift", "f10")

edit.Press_Hotkey("ctrl", "a")

pynput is installed automatically with uiautomation-new. All named pynput keys can be passed as lowercase strings, including f1 through f24, print_screen, caps_lock, and media keys.

Screenshot, Highlight, and UI Tree Export

New debugging and reporting helpers:

control.Screenshot("control.png")
control.Screenshot("control_offscreen.png", offscreen=True)
control.Highlight()

data = control.ToDict()
control.ExportTreeToJson("tree.json", maxDepth=5)
control.ExportTreeToMarkdown("tree.md", maxDepth=5)

offscreen=True uses Windows PrintWindow to capture native windows even when they are covered by other top-level windows. Some hardware-accelerated surfaces such as video, games, WebView, and parts of modern browsers may still return blank or stale content because the target application must support offscreen painting.

Module-level tree export helpers:

  • ControlToDict()
  • ControlTreeToDict()
  • ExportControlTreeToJson()
  • ControlTreeToMarkdown()
  • ExportControlTreeToMarkdown()

Practical Browser Automation Example

The following example shows how the newer helpers can be used together in a browser-like window:

import uiautomation as auto

def get_browser_object(shop_name, class_name=None):
    # 获取对象
    doc_windows = find_windows_by_class_and_title(class_name, shop_name)
    if len(doc_windows) == 0:
        return False
    window = auto.ControlFromHandle(doc_windows[0])
    if window.Exists():
        print(f"成功获取窗口控制对象2: {window2}")
    return window

def get_all_window_controls():
    # 1. Safely enumerate visible top-level windows through Win32.
    top_windows = auto.GetTopLevelControls()
    print(f"共找到 {len(top_windows)} 个顶级窗口:\n")
    # 2. 遍历并打印每个窗口的信息
    for i, window in enumerate(top_windows):
        try:
            print(f"【窗口 {i + 1}】")
            print(f"  窗口标题: {window.Name}")
            print(f"  类名: {window.ClassName}")
            print(f"  句柄: {window.NativeWindowHandle}")
            print(f"  控件类型: {window.ControlType}")
            print(f"  是否可见: {window.Exists()}")
            print("-" * 50)
            print(window.GetChildren())
        except Exception as e:
            print(f"读取窗口失败: {e}")

auto.SetDpiAwareness()

window = auto.WindowControl(ClassName="Chrome_WidgetWin_1", RegexName=".*Google Chrome")

# Enumerate tabs and read their URLs.
tabs = window.FindAll(auto.ControlType.TabItemControl, maxDepth=10)
for tab in tabs:
    url = tab.GetTabUrl(window=window)
    print(tab.Name, url)

# Read all tab URLs in one call. This switches tabs and restores the original tab.
tab_urls = window.GetAllTabUrls()
for item in tab_urls:
    print(item["index"], item["name"], item["url"])

# Switch tabs by title or by URL substring.
window.SwitchToTabByName("PyPI")
window.SwitchToTabByUrl("pypi.org")

# Read the active tab URL from the browser address bar.
url = window.GetCurrentUrl()
print(url)

# Use pynput-backed hotkeys on a focused control.
window = auto.WindowControl(ClassName='Chrome_WidgetWin_1', RegexName='.*Google Chrome')
control = window.EditControl(RegexName='地址和搜索栏')
control.Press_Hotkey("ctrl", "a")
control.Press_Hotkey("backspace")
control.SendKeys('https://pypi.org/project/uiautomation-new/')
# control.Press_Hotkey("ctrl", "v")
control.Press_Hotkey("enter")

# Safely invoke a covered button; SafeClick prefers InvokePattern and falls back to mouse click.
window2 = get_browser_object('万达云')
# 控件隐藏也能点击到
control = window2.ButtonControl(RegexName='智能分流') #
print(control.IsVisibleOnScreen())
control.SafeClick()

# Capture a control even when another top-level window covers it.
address_bar.Screenshot("address_bar.png", offscreen=True)
address_bar.Highlight()

# Export UI tree data for debugging or reports.
window.ExportTreeToJson("tree.json", maxDepth=5)
window.ExportTreeToMarkdown("tree.md", maxDepth=5)

You can also start from a native window handle, which is useful when a window is easier to locate with Win32 APIs:

import win32gui
import uiautomation as auto

def find_windows_by_class_and_title(class_name=None, title_keyword=""):
    handles = []

    def callback(hwnd, _):
        current_class = win32gui.GetClassName(hwnd)
        title = win32gui.GetWindowText(hwnd)
        if class_name:
            matched = current_class == class_name and title_keyword in title
        else:
            matched = title_keyword in title
        if matched:
            handles.append(hwnd)
        return True

    win32gui.EnumWindows(callback, 0)
    return handles

handles = find_windows_by_class_and_title("Chrome_WidgetWin_1", "Google Chrome")
if handles:
    window = auto.ControlFromHandle(handles[0])
    print(window.Name, window.ClassName, window.NativeWindowHandle)

Image Template Matching

match = auto.FindImageOnScreen("button.png", confidence=0.9)
match = auto.WaitImageAppear("button.png", timeout=10)
auto.ClickImage("button.png", timeout=10)

# The template can also be encoded image bytes.
with open("button.png", "rb") as f:
    image_bytes = f.read()
match = auto.FindImageOnScreen(image_bytes, confidence=0.9)

# A captured Bitmap can be converted to PNG bytes directly.
root = auto.GetRootControl()
bitmap = root.ToBitmap(100, 200, 500, 300)
if bitmap:
    with bitmap:
        png_bytes = bitmap.Bytes
        # Use bitmap.ToBytes('jpg') or bitmap.ToBytes('BGRA') for other formats.

# DPI/resolution-tolerant multi-scale matching.
match = auto.FindImageOnScreen(
    "button.png",
    confidence=0.8,
    multiScale=True,
    maxScaleDiff=0.3,
    scaleStep=0.05,
)
auto.ClickImage("button.png", timeout=10, multiScale=True)

# Reuse robust matching settings across multiple operations.
matcher = auto.RobustTemplateMatch(threshold=0.7, maxScaleDiff=0.3)
matcher.click_by_template("button.png", timeout=10)


def image_click():
    path = r'C:\Users\YHCX\Desktop\旧\ScreenShot_2026-07-27_165832_374.png'
    match_result = auto.FindImageOnScreen(path, confidence=0.85)
    print(match_result)
    
    # 1. 解包数据
    x, y, w, h, confidence = match_result

    # 2. 计算中心点
    center_x = x + w // 2  # 780
    center_y = y + h // 2  # 161
    
    # 3. 移动鼠标
    # 方法 A: 如果你的 auto 库有 MoveTo 方法 (根据你之前的代码风格推测)
    auto.MoveTo(center_x, center_y)
    
    # 4. 移动鼠标到坐标点进行点击
    auto.Click(center_x, center_y)

    # 5. 静默点击这个坐标
    auto.SafeClick(center_x, center_y)

    # 6. 传入图片路径 -> 找到图片所在坐标 -> 进行点击
    auto.ClickImage(path)

Multi-scale matching searches around the current Windows DPI scale and also tests the original template size. It returns screen coordinates correctly for restricted regions and virtual desktops. This is useful for custom-rendered UI that does not expose reliable UIAutomation controls. OpenCV and NumPy are installed automatically with the package.

RapidOCR Text Position Clicks

OCR helpers are useful for UIA-invisible text, self-drawn controls, web pages, Canvas, and image-like buttons. Install the OCR extra when needed:

pip install "uiautomation-new[ocr]"

Create a reusable RapidOCR instance. params are initialization parameters for RapidOCR(params=...); they are merged with the library defaults, which use ONNXRuntime for detection, classification, and recognition:

ocr = auto.CreateRapidOCR(
    params={
        "Global.text_score": 0.8,
        "Global.log_level": "warning",
        "Rec.lang_type": "en",
    },
)

Call rapidOCR_img() directly when the raw RapidOCR result_json is needed. Its params argument configures engine creation, while keyword arguments are passed to the recognition call:

result_json = auto.rapidOCR_img(
    "control.png",
    params={"Global.text_score": 0.8},
    use_cls=False,
    text_score=0.8,
)
print(result_json)
# [{"box": [[10, 20], [120, 20], [120, 50], [10, 50]],
#   "txt": "Login", "score": 0.998}]

Recognize normalized text positions from an image file:

positions = auto.OCRImageTextPositions(
    "control.png",
    text="Login",
    minScore=0.8,
    ocr=ocr,
    predictArgs={"use_cls": False, "text_score": 0.8},
)

for item in positions:
    print(item["text"], item["score"], item["box"], item["screen_center"])

One-step OCR text click on a control:

window = auto.WindowControl(ClassName="Chrome_WidgetWin_1", RegexName=".*Google Chrome")

match = auto.OCRClickText("登录", window, minScore=0.8, offscreen=True)
print(match["text"], match["screen_center"]) if match else print("not found")

# Equivalent control method.
match = window.OCRClickText("登录", minScore=0.8, offscreen=True)

For tuning or reuse:

def show_window(handle):
    """ cmdShow 值及效果
    SW_HIDE     0   隐藏窗口(最常用,让程序在后台运行不显示界面)
    SW_SHOW     5   显示窗口(如果之前被隐藏了,就恢复显示)
    SW_MINIMIZE 6   最小化窗口到任务栏
    SW_MAXIMIZE 3   最大化窗口(全屏)
    SW_RESTORE  9   还原窗口(从最小化或最大化恢复到之前的大小)
    :return:
    """
    # 展示最初的桌面
    # auto.ShowDesktop()
    # 显示窗口的状态
    # auto.ShowWindow(handle, cmdShow)

def text_click():
    # 使用文本进行点击
    window = auto.WindowControl(ClassName='Chrome_WidgetWin_1', RegexName='.*Google Chrome')
    control = window.DocumentControl(Name='搜索 - Microsoft 必应')
    match = auto.OCRClickText(
        "国际版",
        control,
        minScore=0.8,
        # 如果控件/窗口被其他顶层窗口遮挡
        offscreen=True,
        ocrArgs={
            "Global.text_score": 0.8,
            "Rec.lang_type": "ch",
        },
        # RapidOCR 本次识别参数
        predictArgs={
            "use_cls": False,
            "text_score": 0.8,
        },
        # SafeClick 开启由Click进行兜底点击
        safeClickArgs={
            "fallbackToClick": True,
        },
    )
    print(match)

OCRClickText() screenshots the control through Screenshot() / ToBitmap(), runs RapidOCR, matches text, converts OCR image coordinates to screen coordinates, and clicks with SafeClick(). center is the OCR center inside the screenshot image; screen_center is center plus the screenshot's actual desktop offset, so it can be used for Windows screen clicks. ocrArgs={...} are passed to RapidOCR(params=...), or create and reuse an OCR instance with CreateRapidOCR(params=...). predictArgs={...} are passed to the RapidOCR recognition call, for example use_cls, use_det, use_rec, text_score, box_thresh, and unclip_ratio. OCR results are normalized to dictionaries with text, score, box, center, screen_box, and screen_center.

Metadata

Release files for uiautomation-new 0.1.5

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for uiautomation-new 0.1.5
File Size Uploaded
uiautomation_new-0.1.5.tar.gz 317.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for uiautomation-new 0.1.5
File Interpreter ABI Platform
uiautomation_new-0.1.5-cp312-cp312-win_amd64.whl CPython 3.12 CPython 3.12 Windows x86-64 Details

Total release size: 785.9 kB

Release files / uiautomation_new-0.1.5.tar.gz

Download URL uiautomation_new-0.1.5.tar.gz
Size 317.4 kB
Tags Source
SHA-256 checksum
How to use checksums
ea53be884b0ce5d1788151d35d88d57617e53e299159b9186a6eaa8b7eb0eb09
BLAKE2b-256 checksum
How to use checksums
b048be1026ec5a83b533eebec34df4deaf808b4ddb1ab0f73b840ee909cadf51
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.2

Release files / uiautomation_new-0.1.5-cp312-cp312-win_amd64.whl

Download URL uiautomation_new-0.1.5-cp312-cp312-win_amd64.whl
Size 468.4 kB
Tags CPython 3.12 Windows x86-64
SHA-256 checksum
How to use checksums
4878ab64ca7902e6079c8cec6c5729f22146267f3eb753315ea4be9cce14eaeb
BLAKE2b-256 checksum
How to use checksums
9fea46235536e0f0dc98e61606cbc2acc92864836f8ddb66b1cd77fb7ccaf852
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.2

Release history Release notifications | RSS feed

This release

0.1.5 This release

2 release files

0.1.0

2 release files

0.0.10

2 release files

0.0.9

2 release files

0.0.8

2 release files

0.0.7

2 release files

0.0.6

2 release files

0.0.5

2 release files

0.0.4

2 release files

0.0.3

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page