uiautomation_new - Refactored Python UIAutomation for Windows
uiautomation-new 后续不再更新,转到 winuia-auto
uiautomation_new is a refactored Windows UI Automation package based on the classic uiautomation API. It keeps the familiar import style and most legacy function names while adding a cleaner internal architecture, modern input helpers, safer clicking, tab-window utilities, visual matching helpers, and PyPI-ready packaging metadata.
This package is designed for automating Windows desktop applications that expose Microsoft UIAutomation providers, including Win32, WinForms, WPF, Modern UI, Qt-based apps, browser windows, and Electron-based apps when accessibility support is enabled.
Installation
pip install uiautomation-new
The default installation includes comtypes, pywin32, Pillow, numpy,
opencv-python, and pynput. Image matching and hotkey helpers work without
installing feature extras. This project does not depend on pyautogui.
Development tools remain optional:
pip install "uiautomation-new[dev]"
Compatibility
The public compatibility layer is intentionally preserved:
import uiautomation as auto
window = auto.GetForegroundControl()
button = window.ButtonControl(Name="OK")
button.Click()
Most old names such as Click, SendKeys, ControlType, PatternId, PropertyId, WindowControl, ButtonControl, and EditControl remain available from the top-level package. New code can also use the added helper APIs described below.
What Changed in This Refactor
Modular Architecture
The old project was centered around one very large uiautomation.py file. This refactor starts splitting responsibilities into focused modules:
uiautomation/enums.py: UIAutomation and Win32 integer enums such asControlType,PatternId,PropertyId,Keys,MouseEventFlag, andKeyboardEventFlag.uiautomation/structures.py:Rectand ctypes input structures such asINPUT,MOUSEINPUT,KEYBDINPUT, andHARDWAREINPUT.uiautomation/inputs/core.py: low-levelSendInput,MouseInput,KeyboardInput, and virtual-key scan-code helpers.uiautomation/inputs/keyboard.py:pynputhotkey helpers.uiautomation/win32_api.py: pywin32-first, ctypes-fallback wrappers for selected Win32 APIs.uiautomation/client.py: thread-safe singleton helpers and package DLL path handling.
The original uiautomation.py remains as a compatibility aggregation layer while larger areas such as controls, patterns, bitmap handling, and logging are migrated progressively.
Modernized Constants and Rect
- Classic constant classes have been moved to
enum.IntEnumwhile keeping int-compatible behavior. Rectis now a dataclass-compatible helper instructures.py.- Empty rectangle detection now treats zero or negative width/height as empty.
Modern Input Simulation
- Mouse clicks, mouse button down/up, mouse wheel, single-key input, and
SendKeysnow useSendInputinternally. - Deprecated
mouse_eventandkeybd_eventnames are still available as compatibility wrappers. SendInputbatches input events instead of sending them one by one.
Human-like Input Simulation
The standard input APIs can opt into human-like timing and movement with
humanLike=True. This keeps the existing API surface while adding smoother
mouse paths, randomized click timing, randomized typing intervals, and more
natural wheel scrolling when explicitly requested.
import uiautomation as auto
# Move with a cubic Bezier path by default.
auto.MoveTo(500, 300, humanLike=True, duration=0.8)
# Click with human-like movement, pre-click pause, and mouse-down hold time.
auto.Click(500, 300, humanLike=True, duration=0.8)
control.Click(humanLike=True, duration=0.8, jitter=4)
# Type with variable intervals. duration controls how long the mouse takes to
# reach the control before typing. errorRate is optional and defaults to 0.
control.SendKeys(
"hello from uiautomation",
humanLike=True,
duration=0.8,
intervalRange=(0.05, 0.18),
errorRate=0.01,
)
# Scroll with non-uniform wheel intervals and a controlled mouse-arrival time.
control.WheelDown(wheelTimes=5, humanLike=True, duration=0.8, intervalRange=(0.1, 0.3))
MoveTo() accepts curveFunc for advanced callers that need a custom movement
curve. The function receives (startX, startY, targetX, targetY, t) and returns
(x, y) for the current progress value:
def linear_curve(start_x, start_y, target_x, target_y, t):
return (
start_x + (target_x - start_x) * t,
start_y + (target_y - start_y) * t,
)
auto.MoveTo(500, 300, humanLike=True, duration=0.8, curveFunc=linear_curve)
Use duration=seconds when the mouse should take a fixed amount of time to
reach the target. Use durationRange=(minSeconds, maxSeconds) when the arrival
time should be randomized. duration takes priority when both are supplied.
SafeClick(humanLike=True) skips silent UIA invoke/post-message paths and uses
the physical mouse click path. Use these options only for normal desktop UI
automation where realistic input timing is desired; they are not intended for
bypassing access controls or verification challenges.
DPI and Multi-Monitor Improvements
New helpers improve high-DPI and multi-monitor coordinate behavior:
import uiautomation as auto
auto.SetDpiAwareness()
rect = auto.GetVirtualScreenRect()
inside = auto.IsPointInVirtualScreen(x, y)
SetDpiAwareness() prefers PerMonitorV2 awareness and automatically falls back to older Windows DPI APIs.
Safer Clicks
Controls now support visibility and clickability checks:
control.IsVisibleOnScreen()
control.IsClickable()
control.SafeClick()
auto.SafeClick(x, y)
control.SafeClick() now prefers UIA InvokePattern for center-click style calls. This is more reliable for buttons whose visual center is covered by another layer or whose physical clickable point is inaccurate. If InvokePattern is unavailable or fails, it falls back to the previous mouse-click path with ScrollIntoView and clickability/occlusion checks.
auto.SafeClick(x, y) is a coordinate-based safe click: it gets the control at (x, y), tries silent UIA actions on the hit control and its ancestors, then tries a Win32 PostMessage click, and falls back to auto.Click(x, y) only when those silent attempts fail. PostMessage can only confirm that mouse messages were posted; some modern or hardware-accelerated apps may ignore them.
To force the physical mouse-click path:
control.SafeClick(preferInvoke=False)
auto.SafeClick(x, y, useInvoke=False, useLegacy=False, usePostMessage=False)
Fluent Relative Search
Controls can be filtered by relative position:
import uiautomation as auto
window = auto.ControlFromHandle(handle) # handle为窗口句柄
save = window.ButtonControl(Name="Save").RightOf(cancel_button)
save = window.ButtonControl(Name="Save").right_of(cancel_button)
Available helpers:
LeftOf()/left_of()RightOf()/right_of()Above()/above()Below()/below()AddCompare()
def relative_position():
# 先定位一个已知的参考控件(作为“锚点”)
reference = auto.ButtonControl(Name="智能分流")
# 然后使用相对位置查找目标 A.RightOf(B) 表示“A 在 B 的右边”
target = auto.ButtonControl(Name="全局代理").RightOf(reference)
print(target)
# 或者写成链式:
target1 = reference.LeftOf(auto.TextControl(Name="全局代理")) # 反过来也行?
# # 注意:通常的写法是:目标控件的条件 .RightOf(参考控件)
# 目标(智能分流) 在 参照物(全局代理)的左侧
print(target1)
#目标(快速连接) 在 参照物(智能分流)的上方
# target2 = auto.TextControl(RegexName="快速连接").Above(reference)
# print(target2)
Smart Waiting
Modern wait aliases are available:
control.WaitUntilExists(timeout=10)
control.WaitUntilDisappear(timeout=5)
auto.WaitUntilExists(control, timeout=10)
auto.WaitUntilDisappear(control, timeout=5)
UI Automation Event Listening
EventHub can listen to UI Automation events such as focus changes, window opened events, and property changes. Event delivery depends on COM message pumping; EventHub(..., autoPump=True) starts a background pump thread. Keep callbacks lightweight.
Typical use cases are places where event-driven waiting is cleaner than polling:
- Wait for a login, save, or export dialog to open before finding its controls.
- Wait for a button, menu item, or text box to become enabled before clicking or typing.
- Record focus changes while debugging or building a simple script recorder.
import time
import uiautomation as auto
from uiautomation import uiautomation as auto_core
hub = auto.EventHub(auto_core._AutomationClient.instance().IUIAutomation)
def on_focus(sender):
control = auto.Control.CreateControlFromElement(sender)
print(control.ControlTypeName, control.Name)
with hub.listen_focus_changed(on_focus):
time.sleep(10)
Property waits can avoid polling when the current value is not ready:
value = auto.WaitForProperty(
button,
auto.PropertyId.IsEnabledProperty,
lambda current: bool(current),
timeout=10,
hub=hub,
)
Recursive FindAll and Tab Utilities
Windows with tab controls can be automated with FindAll() and tab helpers:
tabs = window.FindAll(auto.ControlType.TabItemControl, maxDepth=10)
for tab in tabs:
url = auto.GetTabUrl(tab, window=window)
print(tab.Name, url)
current_tab = window.GetCurrentTab()
window.SwitchToTabByName("Settings")
window.SwitchToTabByUrl("example.com")
url = window.GetCurrentUrl()
tab_urls = window.GetAllTabUrls()
Module-level helpers are also available:
tabs = auto.FindAll(window, auto.ControlType.TabItemControl)
tabs = auto.find_control_num(window, auto.ControlType.TabItemControl)
tab = auto.get_current_tab(window)
auto.switch_to_tab_by_name(window, "Settings")
auto.SwitchToTabByUrl(window, "example.com")
url = auto.get_current_url(window)
url = auto.GetTabUrl(tab, window=window)
tab_urls = auto.GetAllTabUrls(window)
FindAll() supports maxDepth to avoid walking extremely large UI trees indefinitely. Current-tab detection first uses SelectionItemPattern.IsSelected, then falls back to keyboard-focus properties. Tab switching prefers SelectionItemPattern.Select() and falls back to Click(). SwitchToTabByUrl() supports substring URL matching by default. GetTabUrl() and GetAllTabUrls() switch tabs to read the active address bar and then restore the originally selected tab when possible.
Chromium Browser Facade
Version 0.0.7 adds Browser, a high-level facade for a caller-provided Chromium
window. It reuses the existing tab and URL helpers, locates localized address
bars, and waits for a navigation change followed by stable UIA state.
browser = auto.Chrome()
browser.Navigate("https://example.com")
browser.WaitPageLoad(timeout=30)
def on_focus_changed(sender):
print("focused:", sender.Name)
handler = auto.CreateFocusChangedEventHandler(on_focus_changed)
auto.AddFocusChangedEventHandler(handler)
try:
# Browser/UI Automation work...
pass
finally:
auto.RemoveFocusChangedEventHandler(handler)
WaitPageLoad() waits for the URL/title/document UIA state to change and then
remain stable for consecutive samples. It does not guarantee that every
background network request has completed. For automatic cleanup, use the
context-manager form:
with browser.ListenFocusChanged(on_focus_changed):
# Focus events are delivered while this block is active.
pass
Safe Top-Level Window Enumeration
Version 0.1.5 adds GetTopLevelWindowHandles() and GetTopLevelControls().
They enumerate top-level windows through Win32 EnumWindows before optionally
converting each HWND to a UI Automation control. Prefer these helpers when
enumerating desktop windows instead of walking GetRootControl().GetChildren();
individual UIA providers can still fail during ControlFromHandle().
handles = auto.GetTopLevelWindowHandles()
windows = auto.GetTopLevelControls()
Non-UIA Windows
NonUIAWindow is a separate HWND-based facade for windows that do not expose
reliable UI Automation controls. Use WindowControl and other Control
classes for UIA-capable applications; use NonUIAWindow for window-level
automation, screenshots, template matching, OCR, and screen-coordinate input.
window = auto.NonUIAWindow(
title_keyword="My App",
class_name="MyWindowClass",
)
window.activate()
window.click(120, 48) # relative to the window's top-left corner
window.send_keys("hello")
window.screenshot("window.png")
match = window.find_image("save.png", confidence=0.9)
window.click_image("save.png", timeout=5)
Pass hwnd when it is already known; it takes priority over keywords. Without
an HWND, title_keyword and class_name each match by exact value or
substring. A lookup that matches zero or multiple windows raises LookupError.
to_bitmap() tries PrintWindow first and falls back to a Win32/GDI screen
capture when needed. Image matching is automatically limited to the current
window rectangle. click_image() validates the latest window bounds before
clicking the matched center.
close() sends WM_CLOSE by default, so an application can close its window
while keeping background processes alive. Pass terminateProcess=True to
force-stop the owning process; add killProcessTree=True to stop its child
processes too. Forced termination does not save unsaved application data.
window.close()
window.close(terminateProcess=True, killProcessTree=True)
# Stop a known application root process and all of its descendants.
auto.TerminateProcessTree(root_process_id)
If an application destroys its HWND after hide() or set_topmost(), those
methods return False instead of reporting a successful state change.
OCR is optional:
pip install "uiautomation-new[ocr]"
match = window.find_text("OK", minScore=0.8, offscreen=True)
window.click_text("OK", minScore=0.8, offscreen=True)
If a window later exposes usable UIA controls, conversion is explicit and optional:
uia_window = window.to_uia_control()
if uia_window:
uia_window.ButtonControl(Name="OK").Click()
import uiautomation as auto
window = auto.WindowControl(
ClassName="Chrome_WidgetWin_1",
RegexName=".*Google Chrome",
)
browser = auto.Browser(window)
browser.navigate("https://example.com", timeout=30)
print(browser.get_current_url())
print(browser.get_current_tab().Name)
browser.select_tab("Example Domain")
browser.select_tab_by_url("example.com")
for tab in browser.get_all_tab_urls():
print(tab["index"], tab["name"], tab["url"])
Browser currently targets Chromium windows such as Chrome and Edge. Pass the
window explicitly to avoid attaching to an unintended browser instance. Page
loading is inferred from URL, title, and document state; it is not a guarantee
that a web application has completed every background request. Renderer content
still requires the browser accessibility configuration described below.
P1 also provides factories that filter the matching Chromium window by process name, plus browser-standard history commands:
chrome = auto.Chrome(title="Example Domain")
edge = auto.Edge(title="Example Domain")
chrome.back()
chrome.forward()
chrome.reload()
document = chrome.get_current_document()
The factories raise LookupError when the requested executable is not running
or no window title matches. get_current_document() prefers a visible
DocumentControl; renderer content requires --force-renderer-accessibility.
When no browser history entry exists, back() and forward() return False
after the default one-second no_change_timeout rather than waiting for the
full navigation timeout. reload() waits for stable UIA state even when the
URL does not change; use wait=False when only issuing the shortcut matters.
The real browser integration test starts an isolated temporary profile and is disabled by default:
$env:UIAUTOMATION_RUN_BROWSER_INTEGRATION = "1"
$env:UIAUTOMATION_BROWSER = "chrome" # or "edge" or "firefox"
python tests\test_browser_integration.py
Firefox Profile and Page Extraction
Firefox can use the same facade after locating its top-level window and passing
FIREFOX_PROFILE. Its address-bar profile supports the standard labels and
the variable search-engine label used by Firefox.
firefox_window = auto.WindowControl(
ClassName="MozillaWindowClass",
SubName="Example Domain",
)
firefox = auto.Browser(firefox_window, profile=auto.FIREFOX_PROFILE)
Extract UIA content from the active document:
links = browser.get_all_links()
for link in links:
print(link.Name)
text = browser.get_all_text()
print(text)
get_all_text() prefers TextPattern and falls back to visible TextControl
names. Chromium may expose those controls through the window tree rather than
as direct document children; the facade filters that fallback to controls with
a DocumentControl ancestor. UIA text is not a DOM serialization and may omit
content that the browser accessibility provider does not expose.
Browser Tabs and JavaScript Modes
Use browser-standard shortcuts to create or close tabs. new_tab(url) opens a
URL in the newly active tab. close_tab() closes the active tab by default,
closes one zero-based index, or closes every tab matching title and URL filters.
browser.new_tab()
browser.new_tab("https://example.com")
browser.close_tab()
browser.close_tab(2)
browser.close_tab(
"Example Domain",
url="https://example.com/orders/42",
exact=True,
)
For close_tab(), an integer tab is always treated as the tab index and
ignores url, so it closes only that tab. A non-empty string tab closes all
title matches; add a non-empty url to require both values. Pass tab=""
with a URL to close all URL matches. Matching tabs are closed in descending
index order so index changes after each close do not affect later matches.
Passing both tab="" and url="" raises ValueError.
For explicit page-side automation, the default JavaScript mode enters a
javascript: URL through the address bar:
browser.execute_javascript("document.title = 'Ready'")
Use method="console" when the script must be issued through DevTools. It
focuses the Console prompt and types the script directly, avoiding DevTools
paste protection. If this call opens DevTools, it closes it after execution by
default; set close_console=False to leave it open. A Console that was already
open before the call is never toggled or closed.
browser.execute_javascript(
"document.title = 'Ready'",
method="console",
close_console=True,
)
Set paste=True to use the text clipboard instead. The facade restores the
prior text clipboard value before returning. Some DevTools sessions block
pasting until the user enters allow pasting; pass allow_pasting=True to
attempt that input first:
browser.execute_javascript(
"document.title = 'Ready'",
method="console",
paste=True,
allow_pasting=True,
)
Chrome may intentionally reject an automated allow pasting entry as part of
its self-XSS protection. In that case, type the phrase once in the DevTools
console manually, then call the paste=True mode. Neither mode retrieves a
JavaScript return value or bypasses browser security policy. The caller is
responsible for the script's effects.
pynput Hotkey Helper
auto.Press_Hotkey("ctrl", "a")
auto.Press_Hotkey("backspace")
auto.Press_Hotkey("ctrl", "shift", "a")
auto.Press_Hotkey("shift", "f10")
edit.Press_Hotkey("ctrl", "a")
pynput is installed automatically with uiautomation-new.
All named pynput keys can be passed as lowercase strings, including
f1 through f24, print_screen, caps_lock, and media keys.
Screenshot, Highlight, and UI Tree Export
New debugging and reporting helpers:
control.Screenshot("control.png")
control.Screenshot("control_offscreen.png", offscreen=True)
control.Highlight()
data = control.ToDict()
control.ExportTreeToJson("tree.json", maxDepth=5)
control.ExportTreeToMarkdown("tree.md", maxDepth=5)
offscreen=True uses Windows PrintWindow to capture native windows even when they are covered by other top-level windows. Some hardware-accelerated surfaces such as video, games, WebView, and parts of modern browsers may still return blank or stale content because the target application must support offscreen painting.
Module-level tree export helpers:
ControlToDict()ControlTreeToDict()ExportControlTreeToJson()ControlTreeToMarkdown()ExportControlTreeToMarkdown()
Practical Browser Automation Example
The following example shows how the newer helpers can be used together in a browser-like window:
import uiautomation as auto
def get_browser_object(shop_name, class_name=None):
# 获取对象
doc_windows = find_windows_by_class_and_title(class_name, shop_name)
if len(doc_windows) == 0:
return False
window = auto.ControlFromHandle(doc_windows[0])
if window.Exists():
print(f"成功获取窗口控制对象2: {window2}")
return window
def get_all_window_controls():
# 1. Safely enumerate visible top-level windows through Win32.
top_windows = auto.GetTopLevelControls()
print(f"共找到 {len(top_windows)} 个顶级窗口:\n")
# 2. 遍历并打印每个窗口的信息
for i, window in enumerate(top_windows):
try:
print(f"【窗口 {i + 1}】")
print(f" 窗口标题: {window.Name}")
print(f" 类名: {window.ClassName}")
print(f" 句柄: {window.NativeWindowHandle}")
print(f" 控件类型: {window.ControlType}")
print(f" 是否可见: {window.Exists()}")
print("-" * 50)
print(window.GetChildren())
except Exception as e:
print(f"读取窗口失败: {e}")
auto.SetDpiAwareness()
window = auto.WindowControl(ClassName="Chrome_WidgetWin_1", RegexName=".*Google Chrome")
# Enumerate tabs and read their URLs.
tabs = window.FindAll(auto.ControlType.TabItemControl, maxDepth=10)
for tab in tabs:
url = tab.GetTabUrl(window=window)
print(tab.Name, url)
# Read all tab URLs in one call. This switches tabs and restores the original tab.
tab_urls = window.GetAllTabUrls()
for item in tab_urls:
print(item["index"], item["name"], item["url"])
# Switch tabs by title or by URL substring.
window.SwitchToTabByName("PyPI")
window.SwitchToTabByUrl("pypi.org")
# Read the active tab URL from the browser address bar.
url = window.GetCurrentUrl()
print(url)
# Use pynput-backed hotkeys on a focused control.
window = auto.WindowControl(ClassName='Chrome_WidgetWin_1', RegexName='.*Google Chrome')
control = window.EditControl(RegexName='地址和搜索栏')
control.Press_Hotkey("ctrl", "a")
control.Press_Hotkey("backspace")
control.SendKeys('https://pypi.org/project/uiautomation-new/')
# control.Press_Hotkey("ctrl", "v")
control.Press_Hotkey("enter")
# Safely invoke a covered button; SafeClick prefers InvokePattern and falls back to mouse click.
window2 = get_browser_object('万达云')
# 控件隐藏也能点击到
control = window2.ButtonControl(RegexName='智能分流') #
print(control.IsVisibleOnScreen())
control.SafeClick()
# Capture a control even when another top-level window covers it.
address_bar.Screenshot("address_bar.png", offscreen=True)
address_bar.Highlight()
# Export UI tree data for debugging or reports.
window.ExportTreeToJson("tree.json", maxDepth=5)
window.ExportTreeToMarkdown("tree.md", maxDepth=5)
You can also start from a native window handle, which is useful when a window is easier to locate with Win32 APIs:
import win32gui
import uiautomation as auto
def find_windows_by_class_and_title(class_name=None, title_keyword=""):
handles = []
def callback(hwnd, _):
current_class = win32gui.GetClassName(hwnd)
title = win32gui.GetWindowText(hwnd)
if class_name:
matched = current_class == class_name and title_keyword in title
else:
matched = title_keyword in title
if matched:
handles.append(hwnd)
return True
win32gui.EnumWindows(callback, 0)
return handles
handles = find_windows_by_class_and_title("Chrome_WidgetWin_1", "Google Chrome")
if handles:
window = auto.ControlFromHandle(handles[0])
print(window.Name, window.ClassName, window.NativeWindowHandle)
Image Template Matching
match = auto.FindImageOnScreen("button.png", confidence=0.9)
match = auto.WaitImageAppear("button.png", timeout=10)
auto.ClickImage("button.png", timeout=10)
# The template can also be encoded image bytes.
with open("button.png", "rb") as f:
image_bytes = f.read()
match = auto.FindImageOnScreen(image_bytes, confidence=0.9)
# A captured Bitmap can be converted to PNG bytes directly.
root = auto.GetRootControl()
bitmap = root.ToBitmap(100, 200, 500, 300)
if bitmap:
with bitmap:
png_bytes = bitmap.Bytes
# Use bitmap.ToBytes('jpg') or bitmap.ToBytes('BGRA') for other formats.
# DPI/resolution-tolerant multi-scale matching.
match = auto.FindImageOnScreen(
"button.png",
confidence=0.8,
multiScale=True,
maxScaleDiff=0.3,
scaleStep=0.05,
)
auto.ClickImage("button.png", timeout=10, multiScale=True)
# Reuse robust matching settings across multiple operations.
matcher = auto.RobustTemplateMatch(threshold=0.7, maxScaleDiff=0.3)
matcher.click_by_template("button.png", timeout=10)
def image_click():
path = r'C:\Users\YHCX\Desktop\旧\ScreenShot_2026-07-27_165832_374.png'
match_result = auto.FindImageOnScreen(path, confidence=0.85)
print(match_result)
# 1. 解包数据
x, y, w, h, confidence = match_result
# 2. 计算中心点
center_x = x + w // 2 # 780
center_y = y + h // 2 # 161
# 3. 移动鼠标
# 方法 A: 如果你的 auto 库有 MoveTo 方法 (根据你之前的代码风格推测)
auto.MoveTo(center_x, center_y)
# 4. 移动鼠标到坐标点进行点击
auto.Click(center_x, center_y)
# 5. 静默点击这个坐标
auto.SafeClick(center_x, center_y)
# 6. 传入图片路径 -> 找到图片所在坐标 -> 进行点击
auto.ClickImage(path)
Multi-scale matching searches around the current Windows DPI scale and also tests the original template size. It returns screen coordinates correctly for restricted regions and virtual desktops. This is useful for custom-rendered UI that does not expose reliable UIAutomation controls. OpenCV and NumPy are installed automatically with the package.
RapidOCR Text Position Clicks
OCR helpers are useful for UIA-invisible text, self-drawn controls, web pages, Canvas, and image-like buttons. Install the OCR extra when needed:
pip install "uiautomation-new[ocr]"
Create a reusable RapidOCR instance. params are initialization parameters for
RapidOCR(params=...); they are merged with the library defaults, which use
ONNXRuntime for detection, classification, and recognition:
ocr = auto.CreateRapidOCR(
params={
"Global.text_score": 0.8,
"Global.log_level": "warning",
"Rec.lang_type": "en",
},
)
Call rapidOCR_img() directly when the raw RapidOCR result_json is needed.
Its params argument configures engine creation, while keyword arguments are
passed to the recognition call:
result_json = auto.rapidOCR_img(
"control.png",
params={"Global.text_score": 0.8},
use_cls=False,
text_score=0.8,
)
print(result_json)
# [{"box": [[10, 20], [120, 20], [120, 50], [10, 50]],
# "txt": "Login", "score": 0.998}]
Recognize normalized text positions from an image file:
positions = auto.OCRImageTextPositions(
"control.png",
text="Login",
minScore=0.8,
ocr=ocr,
predictArgs={"use_cls": False, "text_score": 0.8},
)
for item in positions:
print(item["text"], item["score"], item["box"], item["screen_center"])
One-step OCR text click on a control:
window = auto.WindowControl(ClassName="Chrome_WidgetWin_1", RegexName=".*Google Chrome")
match = auto.OCRClickText("登录", window, minScore=0.8, offscreen=True)
print(match["text"], match["screen_center"]) if match else print("not found")
# Equivalent control method.
match = window.OCRClickText("登录", minScore=0.8, offscreen=True)
For tuning or reuse:
def show_window(handle):
""" cmdShow 值及效果
SW_HIDE 0 隐藏窗口(最常用,让程序在后台运行不显示界面)
SW_SHOW 5 显示窗口(如果之前被隐藏了,就恢复显示)
SW_MINIMIZE 6 最小化窗口到任务栏
SW_MAXIMIZE 3 最大化窗口(全屏)
SW_RESTORE 9 还原窗口(从最小化或最大化恢复到之前的大小)
:return:
"""
# 展示最初的桌面
# auto.ShowDesktop()
# 显示窗口的状态
# auto.ShowWindow(handle, cmdShow)
def text_click():
# 使用文本进行点击
window = auto.WindowControl(ClassName='Chrome_WidgetWin_1', RegexName='.*Google Chrome')
control = window.DocumentControl(Name='搜索 - Microsoft 必应')
match = auto.OCRClickText(
"国际版",
control,
minScore=0.8,
# 如果控件/窗口被其他顶层窗口遮挡
offscreen=True,
ocrArgs={
"Global.text_score": 0.8,
"Rec.lang_type": "ch",
},
# RapidOCR 本次识别参数
predictArgs={
"use_cls": False,
"text_score": 0.8,
},
# SafeClick 开启由Click进行兜底点击
safeClickArgs={
"fallbackToClick": True,
},
)
print(match)
OCRClickText() screenshots the control through Screenshot() / ToBitmap(), runs RapidOCR, matches text, converts OCR image coordinates to screen coordinates, and clicks with SafeClick(). center is the OCR center inside the screenshot image; screen_center is center plus the screenshot's actual desktop offset, so it can be used for Windows screen clicks. ocrArgs={...} are passed to RapidOCR(params=...), or create and reuse an OCR instance with CreateRapidOCR(params=...). predictArgs={...} are passed to the RapidOCR recognition call, for example use_cls, use_det, use_rec, text_score, box_thresh, and unclip_ratio. OCR results are normalized to dictionaries with text, score, box, center, screen_box, and screen_center.
Metadata
Release files for uiautomation-new 0.1.5
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| uiautomation_new-0.1.5.tar.gz | 317.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| uiautomation_new-0.1.5-cp312-cp312-win_amd64.whl | CPython 3.12 | CPython 3.12 | Windows x86-64 | Details |
Total release size: 785.9 kB
Release files / uiautomation_new-0.1.5.tar.gz
| Download URL | uiautomation_new-0.1.5.tar.gz |
|---|---|
| Size | 317.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
ea53be884b0ce5d1788151d35d88d57617e53e299159b9186a6eaa8b7eb0eb09
|
|
BLAKE2b-256 checksum How to use checksums |
b048be1026ec5a83b533eebec34df4deaf808b4ddb1ab0f73b840ee909cadf51
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.2
|
Release files / uiautomation_new-0.1.5-cp312-cp312-win_amd64.whl
| Download URL | uiautomation_new-0.1.5-cp312-cp312-win_amd64.whl |
|---|---|
| Size | 468.4 kB |
| Tags | CPython 3.12 Windows x86-64 |
|
SHA-256 checksum How to use checksums |
4878ab64ca7902e6079c8cec6c5729f22146267f3eb753315ea4be9cce14eaeb
|
|
BLAKE2b-256 checksum How to use checksums |
9fea46235536e0f0dc98e61606cbc2acc92864836f8ddb66b1cd77fb7ccaf852
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.2
|