pyvizion
pyvizion automatiza qualquer programa pela tela, como uma pessoa faria: encontra imagens (OpenCV, várias escalas) e textos (OCR com Tesseract), clica, digita e espera — com backtrack automático. Feito para não falhar em sistemas legados: ERPs, Oracle Forms, Delphi/VB6, Java Swing, Citrix/RDP e emuladores de terminal.
from pyvizion import Vizion
vz = Vizion()
vz.click_image("button1.png", backtrack=True)
vz.click_text("Save", backtrack=True) # se falhar, refaz o button1 e tenta de novo
vz.click_text("Confirm", backtrack=True) # se falhar, refaz o Save e tenta de novo
📦 Instalação
pip install pyvizion
Uma instalação só, com tudo incluído e já no modo mais rápido.
Tesseract (só para funções de texto — click_text, find_text, read_text):
- Windows: https://github.com/UB-Mannheim/tesseract/wiki — instale em
C:\Program Files\Tesseract-OCR\e marque Portuguese. - Linux:
sudo apt-get install tesseract-ocr tesseract-ocr-por - macOS:
brew install tesseract tesseract-lang
python -m pyvizion doctor # confere dependências, Tesseract, idiomas e monitores
🎯 Uso rápido
Backtrack entre métodos
vz = Vizion()
vz.click_image("menu.png", backtrack=True)
vz.click_text("Relatórios", backtrack=True) # falhou? reexecuta o anterior e tenta de novo
Sessão de backtrack
vz.start_task_session()
vz.click_image("button1.png", backtrack=True)
vz.click_text("Clientes", backtrack=True)
vz.click_text("Novo", backtrack=True)
successful, total = vz.end_task_session()
print(f"Sucesso: {successful}/{total}")
Lista de tarefas
from pyvizion import execute_tasks
tasks = [
{"image": "button.png", "region": (100, 100, 200, 50), "confidence": 0.9,
"specific": False, "backtrack": True, "delay": 1, "mouse_button": "left"},
{"text": "Login", "region": (50, 50, 300, 100), "char_type": "letters",
"backtrack": True, "sendtext": "usuario123{tab}senha{enter}"},
{"type": "relative_image", "anchor_image": "warning_icon.png",
"target_image": "ok_button.png", "max_distance": 200},
{"type": "click", "x": 500, "y": 300, "mouse_button": "right"},
{"type": "type_text", "text": "Hello World!"},
{"type": "keyboard_command", "command": "Ctrl+S", "delay": 1},
]
execute_tasks(tasks) # ou: Vizion().execute_tasks(tasks)
🔧 Métodos
| Método | O que faz |
|---|---|
click_image(image_path, region, confidence, delay, mouse_button, max_attempts, backtrack, specific, sendtext, show_overlay) |
Encontra e clica numa imagem |
find_image(image_path, region, confidence, max_attempts, backtrack, specific, scales) |
Retorna (x, y, largura, altura) ou None |
click_text(text, region, filter_type, delay, mouse_button, occurrence, backtrack, max_attempts, sendtext, confidence_threshold, show_overlay) |
Encontra e clica num texto |
find_text(text, region, filter_type, confidence_threshold, occurrence, max_attempts, backtrack) |
Caixa do texto ou None |
click_relative_image(anchor_image, target_image, max_distance, confidence, target_region, delay, mouse_button, backtrack, max_attempts) |
Clica no alvo mais próximo de uma âncora |
find_relative_image(...) |
Caixa do alvo mais próximo da âncora |
click_image_near_text(anchor_text, target_image, ...) |
Clica na imagem mais próxima de um texto |
click_coordinates(x, y, delay, mouse_button, backtrack) / click_at(location, ...) |
Clique em coordenadas |
type_text(text, interval, delay, backtrack) |
Digita texto (aceita macros) |
keyboard_command(command, delay, backtrack) |
"Ctrl+S", "F7", "Alt+Tab"... |
wait_for_image(...) / wait_for_text(...) / wait_until_gone(...) |
Esperas |
image_exists(...) / text_exists(...) / find_all_images(...) |
Verificações |
click_any([...]) / find_any([...]) |
Primeiro alvo encontrado dentre alternativas |
read_text(region, single_line) |
Lê o texto de uma área |
focus_window(title) / wait_window(title) / window_region(title) |
Janelas |
press(key, presses) / hotkey(*keys) / scroll(clicks) / drag(start, end) / screenshot(path) |
Utilidades |
execute_tasks(tasks) / execute_with_backtrack_between_tasks(tasks) |
Listas de tarefas |
start_task_session() / end_task_session() |
Sessão de backtrack |
configure_overlay(enabled, color, duration, width) / get_overlay_config() / test_overlay_colors() |
Overlay visual |
list_windows() |
Títulos das janelas abertas |
Parâmetros comuns
region=(x, y, largura, altura)— onde procurar. Use sempre que puder (OCR ~6× mais rápido, sem falsos positivos).vz.window_region("Título")devolve a área de uma janela.specific=True— busca só naregion.specific=False— tenta a região e depois a tela inteira, em escalas 0,8× a 1,2×.mouse_button—"left","right","double","move_to"(só passa o mouse),"middle","triple".filter_type—"letters","numbers"ou"both".occurrence=2— a segunda ocorrência em ordem de leitura.offset=(dx, dy)— clica deslocado do alvo (ex.: o campo à direita do rótulo).timeout=— espera o alvo aparecer por até N segundos.
Os métodos também existem como funções, sem instância: from pyvizion import click_image, find_text, execute_tasks.
📝 Preenchendo campos (sendtext)
sendtext é o texto digitado logo depois do clique. Dentro dele, palavras entre chaves { } são teclas:
vz.click_text("Usuário", offset=(150, 0), sendtext="admin{tab}senha123{enter}")
O que acontece, em ordem:
| # | Trecho | O pyvizion faz |
|---|---|---|
| 1 | (clique) | clica 150 px à direita do rótulo "Usuário" (dentro do campo) |
| 2 | admin |
digita "admin" |
| 3 | {tab} |
aperta Tab → vai para o campo de senha |
| 4 | senha123 |
digita "senha123" |
| 5 | {enter} |
aperta Enter → confirma |
Receitas do dia a dia
| Quero... | sendtext |
|---|---|
| digitar num campo vazio | "12345" |
| substituir o que já está no campo | "{ctrl}a{del}12345" |
| digitar e ir para o próximo campo | "12345{tab}" |
| digitar e confirmar | "12345{enter}" |
| preencher dois campos seguidos | "01/01/2026{tab}31/01/2026" |
| pular campos | "{tab*3}" |
| esperar o sistema reagir | "12345{tab}{wait 1}" |
Prefira um passo por campo — fica mais fácil de ler e, se algo falhar, o log mostra exatamente onde:
vz.click_text("Data Inicial", offset=(150, 0), sendtext="{ctrl}a{del}01/01/2026")
vz.click_text("Data Final", offset=(150, 0), sendtext="{ctrl}a{del}31/01/2026")
vz.keyboard_command("Enter")
Teclas disponíveis
| Macro | Tecla |
|---|---|
{enter} {tab} {esc} {space} |
Enter, Tab, Esc, Espaço |
{del} {backspace} {insert} |
Delete, Backspace, Insert |
{up} {down} {left} {right} |
setas |
{home} {end} {pgup} {pgdn} |
Home, End, Page Up, Page Down |
{f1} … {f12} |
teclas de função |
{ctrl}a {alt}f {shift}x |
modificador + a próxima letra |
{ctrl+shift+s} {alt+f4} |
combinação completa |
{tab*3} {down 5} |
repetição |
{wait 1.5} {wait 300ms} |
pausa |
{{ }} |
as chaves { e } literais |
- Maiúsculas não importam:
{Enter}={enter}. - Chaves com algo que não é tecla são digitadas normalmente:
"valor {total}"digita exatamente isso. - Acentos e
çsaem corretos, e a área de transferência do usuário é restaurada. - Em Citrix/RDP/terminais que não aceitam colar, use
Vizion({"typing_mode": "type"}).
Para apertar só uma tecla, sem texto: vz.keyboard_command("Ctrl+S"), vz.press("tab", 3), vz.hotkey("ctrl", "shift", "s").
📋 Tarefas: todos os tipos
tasks = [
{"type": "focus_window", "title": "Oracle Applications"},
{"type": "wait_image", "image": "tela_principal.png", "timeout": 30},
{"text": "Cliente", "offset": (160, 0), "sendtext": "12345{enter}", "required": True},
{"type": "wait_image", "image": "ampulheta.png", "gone": True, "timeout": 60},
{"type": "relative_image", "anchor_text": "Pedido", "target_image": "lupa.png"},
{"image": "popup_aviso.png", "optional": True, "sendtext": "{enter}"},
{"type": "wait_text", "text": "Registro salvo", "timeout": 15},
{"type": "wait", "seconds": 1},
{"type": "scroll", "clicks": -5},
]
| Tipo | Chaves |
|---|---|
| imagem | image, region, confidence, specific, scales |
| texto | text, region, char_type, occurrence, confidence_threshold |
relative_image |
anchor_image ou anchor_text, target_image, max_distance, target_region |
click |
x, y |
type_text |
text, interval |
keyboard_command |
command |
wait_image / wait_text |
image/text, timeout, gone |
wait · focus_window · scroll |
seconds · title · clicks, x, y |
Chaves de qualquer tarefa: mouse_button, delay, sendtext, offset, backtrack, max_attempts, timeout, optional (falha não conta nem faz backtrack), required (falha interrompe a lista), show_overlay, click_hold.
⚙️ Configuração
vz = Vizion({
"confidence_threshold": 80.0,
"tesseract_lang": "por",
"show_overlay": False,
"image_folders": ["./imagens"],
"save_failure_screenshots": True,
})
vz.config.set("show_overlay", True) # alterar depois
vz = Vizion("pyvizion.json") # ou de um arquivo .json/.yaml
| Chave | Padrão | Descrição |
|---|---|---|
confidence_threshold |
75.0 |
Limiar do OCR (0–100) |
default_confidence |
0.9 |
Confiança padrão de imagens nas tarefas |
tesseract_path / tessdata_path |
auto | Caminhos do Tesseract |
tesseract_lang |
auto (por se instalado) |
Idioma(s): "por", "por+eng" |
image_processing_methods |
"all" |
"all", "balanced", "fast" ou lista de técnicas |
ocr_large_image_methods |
"fast" |
Técnicas usadas na tela inteira |
preprocessing_enabled |
True |
False = OCR só na imagem original |
ocr_upscale / ocr_workers |
2.0 / auto |
Ampliação de áreas pequenas / paralelismo |
ocr_fuzzy / ocr_fuzzy_threshold |
True / 0.8 |
Tolerância a erros do OCR |
grayscale · scales · dpi_scales · edge_fallback |
True · 0,8–1,2 · True · False |
Busca de imagem |
min_confidence · retry_delay |
0.7 · 0.5 |
Tentativas |
overlay_enabled |
True |
Liga/desliga o sistema de overlay |
image_folders |
[] |
Pastas onde procurar imagens |
typing_mode · restore_clipboard · typing_interval |
"paste" · True · 0.02 |
Digitação |
click_hold · move_pause · movement_duration |
0 · 0.15 · 0.1 |
Ritmo do mouse |
failsafe |
True |
Mouse no canto superior esquerdo interrompe |
show_overlay · overlay_color · overlay_duration · overlay_width |
False · red · 1000 · 4 |
Retângulo sobre o alvo antes do clique (depuração) |
save_failure_screenshots · failure_screenshot_dir |
False · pyvizion_failures |
Print a cada falha |
stop_on_failure · max_backtrack_attempts |
False · 2 |
Listas de tarefas |
log_level |
"INFO" |
🛡️ Robustez para sistemas legados
- DPI e vários monitores — modo DPI por monitor, captura de todos os monitores (inclusive coordenadas negativas); a escala de cada monitor (125%, 150%) entra automaticamente na busca.
- OCR tolerante — ignora acentos e maiúsculas, corrige
0/O,1/l,5/S, aceita pequenas diferenças e palavras coladas/quebradas, amplia áreas pequenas, inverte texto claro em fundo escuro, roda em paralelo e usa consenso entre leituras. - Captura rápida — ~5 ms por área.
- Esperas, alternativas (
click_any), âncoras (click_relative_image,click_image_near_text), clique deslocado (offset), leitura de campos (read_text), foco de janela. - Listas seguras —
required,optional,stop_on_failure. - Diagnóstico — prints de falha e
python -m pyvizion doctor.
🛠️ Linha de comando
python -m pyvizion doctor # ambiente
python -m pyvizion screenshot tela.png # print para recortar suas imagens
python -m pyvizion position # x/y do mouse em tempo real (Ctrl+C)
python -m pyvizion pick # marca região (2 cantos), OCR e trechos para colar no código
python -m pyvizion ocr 100 200 300 40 # lê o texto de uma área (se já souber x,y,w,h)
💡 Dicas
- Use
region(ouvz.window_region("Título")) sempre que puder. - Recorte imagens pequenas e únicas (o ícone, não o botão inteiro), em PNG.
- Prefira
wait_until_gone("ampulheta.png")adelayfixo. - Comece com
focus_window("Título"). - Instale o idioma
pordo Tesseract. - Para parar o robô, leve o mouse ao canto superior esquerdo da tela.
🔄 Migrando de bot-vision-suite / visus-desktop
Substituição mínima:
- pip install bot-vision-suite
+ pip install pyvizion
- from bot_vision import BotVision
+ from pyvizion import Vizion
- bot = BotVision()
+ vz = Vizion()
Funções sem instância continuam iguais em espírito (from pyvizion import click_image, execute_tasks, …).
Tabela 1:1 (métodos)
bot-vision-suite / BotVision |
pyvizion / Vizion |
Observação |
|---|---|---|
BotVision() |
Vizion() ou Vizion({...}) / Vizion("config.yaml") |
pyvizion aceita dict, JSON ou YAML |
BotVision(config=dict) |
Vizion(config=dict) + Config |
Chaves parecidas; veja Configuração |
click_image(...) |
click_image(...) |
Compatível |
find_image(...) |
find_image(...) |
Retorno: (x, y, w, h) ou None (mesma ideia) |
find_all_images(...) |
find_all_images(...) |
Compatível |
image_exists(...) |
image_exists(...) |
Compatível |
click_text(...) |
click_text(...) |
Compatível |
find_text(...) |
find_text(...) |
Compatível |
text_exists(...) |
text_exists(...) |
Compatível |
read_text(region, single_line) |
read_text(region, single_line) |
Compatível |
click_relative_image(anchor, target, ...) |
click_relative_image(...) |
Compatível |
find_relative_image(...) |
find_relative_image(...) |
Compatível |
click_image_near_text(anchor_text, target, ...) |
click_image_near_text(...) |
Compatível |
click_coordinates(x, y, ...) |
click_coordinates(x, y, ...) |
Compatível |
click_at(location, ...) |
click_at(location, ...) |
Compatível |
type_text(...) |
type_text(...) |
Macros {enter}, {tab}, {ctrl}a… |
keyboard_command(...) |
keyboard_command(...) |
Compatível |
wait_for_image(...) |
wait_for_image(...) |
Compatível |
wait_for_text(...) |
wait_for_text(...) |
Compatível |
wait_until_gone(...) |
wait_until_gone(image_path=..., text=...) |
Compatível |
click_any([...]) |
click_any([...]) |
pyvizion retorna índice clicado ou None (não só bool) |
find_any([...]) |
find_any([...]) |
Retorno (índice, region) ou (None, None) |
execute_tasks(tasks) |
execute_tasks(tasks) |
Compatível; pyvizion tem ainda optional, required, stop_on_failure |
execute_with_backtrack_between_tasks(...) |
execute_with_backtrack_between_tasks(...) |
Formato {'type', 'params'} |
start_task_session() / end_task_session() |
idem | Retorno (ok, total) |
configure_overlay(...) |
configure_overlay(...) |
Compatível |
get_overlay_config() |
get_overlay_config() |
Compatível |
test_overlay_colors() |
test_overlay_colors() |
Compatível |
focus_window / wait_window / window_region |
idem | Compatível |
list_windows() |
list_windows() |
Compatível |
press / hotkey / scroll / drag / screenshot |
idem em Vizion |
Compatível |
from bot_vision import click_image, … |
from pyvizion import click_image, … |
Mesmo padrão |
python -m bot_vision doctor (se existir) |
python -m pyvizion doctor |
+ screenshot, position, ocr |
Parâmetros que mudam de nome (tarefas / dicts)
| BVS / visus | pyvizion | Notas |
|---|---|---|
"filter_type": "letters" em tarefas |
"char_type": "letters" |
Nos métodos continua filter_type= |
"specific": false (busca tela inteira + escalas) |
"specific": false |
Internamente vira flexible=True na busca |
"anchor_image" + "target_image" |
idem | + "anchor_text" em type: relative_image |
"confidence_threshold" (OCR) |
"confidence_threshold" ou "early_confidence" |
Nos métodos: confidence_threshold= |
O que o pyvizion adiciona (não existe na BVS da mesma forma)
| Recurso | Onde |
|---|---|
Config em JSON/YAML + config.set() com reload |
Vizion("projeto.json") |
Pastas de imagens (image_folders) |
Config |
| Screenshot automático em falha | save_failure_screenshots |
OCR multi-técnica nomeada (fast / balanced / all) |
image_processing_methods |
| DPI por monitor + captura multi-monitor | import + dpi_scales |
| Matching por bordas (tema claro/escuro) | edge_fallback |
| Modo digitação para Citrix/RDP | typing_mode: "type" |
Alvos tipados (Image, Text, Near, AnyOf) |
core/specs.py (API interna/moderna) |
Tarefas optional / required / click_hold |
execute_tasks |
O que ainda falta para drop-in 100% com BotVision
| Item | Situação | Workaround hoje |
|---|---|---|
BotVision alias |
Não exportado | from pyvizion import Vizion as BotVision |
register_image("ok", "ok.png") (visus-desktop) |
Não implementado | Caminho completo ou image_folders: ["./imagens"] |
limpar_texto / clean_text |
Não implementado | Normalizar string antes do OCR manualmente |
Pacote bot_vision |
Nome diferente | Só trocar imports para pyvizion |
| EasyOCR / PyTorch (opcional na BVS) | Não incluso | Só Tesseract (proposital — mais leve) |
Extras AI (openai) da BVS |
Não incluso | Fora do escopo desktop |
| Automação web (Selenium/DOM) | Não incluso | BVS também não tem; use Selenium à parte |
Com a troca de import e de BotVision → Vizion, a maior parte dos scripts BVS roda sem alterar chamadas de click_image, click_text, execute_tasks e backtrack.
🧪 Desenvolvimento
pip install -e .[dev]
pytest
📄 Licença
Software proprietário — veja LICENSE.
Desenvolvido por Josias Azevedo da Silva.
Metadata
Release files for pyvizion 1.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pyvizion-1.0.0.tar.gz | 62.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pyvizion-1.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 129.7 kB
Release files / pyvizion-1.0.0.tar.gz
| Download URL | pyvizion-1.0.0.tar.gz |
|---|---|
| Size | 62.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
43616ee897d2dc583cb7edcdbf1b44b752353aa0605ba6f345c8d5b471aefc07
|
|
BLAKE2b-256 checksum How to use checksums |
f87114e35ee163086c4271dad4a31d10f28b29b97f54a6f4e3ed4d4a6c7fd8e1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.0
|
Release files / pyvizion-1.0.0-py3-none-any.whl
| Download URL | pyvizion-1.0.0-py3-none-any.whl |
|---|---|
| Size | 67.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
bbe79152612ee2861d561e61f845f80e6e8079f65016d9fcf045a393bcbea7b6
|
|
BLAKE2b-256 checksum How to use checksums |
521b2ea73adc348d5d1d0000daa7fdc606abf1c673a4b1430f7835fb2c9a3bf0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.0
|