Skip to main content

Streamlit multimodal chat input component with text, image, and voice support

Project description

Streamlit Multimodal Chat Input

A multimodal chat input component for Streamlit that supports text input, image upload, and voice input.

Note: Voice and image features require HTTPS or localhost environment to function properly.

Features

  • 📝 Text Input: Same usability as st.chat_input
  • 🖼️ Image File Upload: Supports jpg, png, gif, webp
  • 📸 Screenshot Capture: Capture and share screenshots directly
  • 🎤 Voice Input: Web Speech API / OpenAI Whisper API support
  • 🎨 Streamlit Standard Theme: Fully compatible design
  • 🔄 Drag & Drop: File drag and drop support
  • ⌨️ Ctrl+V: Paste images from clipboard
  • ⚙️ Customizable: Rich configuration options

Installation

pip install quadis-chat-input-multimodal

Basic Usage

import streamlit as st
from quadis_chat_input import quadis_chat_input

# Basic usage
result = quadis_chat_input()

if result:
    # Display text
    if result['text']:
        st.write(f"Text: {result['text']}")
    
    # Display uploaded files
    if result['files']:
        for file in result['files']:
            import base64
            base64_data = file['data'].split(',')[1]
            image_bytes = base64.b64decode(base64_data)
            st.image(image_bytes, caption=file['name'])
    
    # Display voice input metadata
    if result.get('audio_metadata'):
        st.write(f"Voice input used: {result['audio_metadata']['used_voice_input']}")

Advanced Usage

Voice Input Features

# Enable voice input
result = quadis_chat_input(
    enable_voice_input=True,
    voice_recognition_method="web_speech",  # or "openai_whisper"
    voice_language="ja-JP",
    max_recording_time=60
)

# Using OpenAI Whisper API
result = quadis_chat_input(
    enable_voice_input=True,
    voice_recognition_method="openai_whisper",
    openai_api_key="sk-your-api-key",
    voice_language="ja-JP"
)

Screenshot Capture

# Enable screenshot capture (enabled by default)
result = quadis_chat_input(
    enable_screenshot=True
)

# Disable screenshot button
result = quadis_chat_input(
    enable_screenshot=False
)

Custom Configuration

result = quadis_chat_input(
    placeholder="Please enter your message...",
    max_chars=500,
    accepted_file_types=["jpg", "png", "gif", "webp"],
    max_file_size_mb=10,
    disabled=False,
    enable_screenshot=True,  # Control screenshot button visibility
    key="custom_chat_input"
)

Chat Usage

import streamlit as st
import base64
from quadis_chat_input import quadis_chat_input

# Page configuration
st.set_page_config(
    page_title="Multimodal Chat Input Demo",
    page_icon="💬",
    layout="wide"
)

st.subheader("💭 Multimodal Chat Input Demo")
st.markdown("Simulate a chat application with voice input and file upload.")

# Manage history in session state
if "chat_history" not in st.session_state:
    st.session_state.chat_history = []

# Input for new messages
chat_result = quadis_chat_input(
    placeholder="Enter chat message...",
    enable_voice_input=True,  # Enable voice input for chat as well
    key="chat_input"
)
if chat_result:
    st.session_state.chat_history.append(chat_result)

# Display chat history
if st.session_state.chat_history:
    for i, message in enumerate(st.session_state.chat_history):
        with st.chat_message("user"):
            if message.get("text"):
                st.write(message["text"])
            
            if message.get("files"):
                for file in message["files"]:
                    try:
                        base64_data = file['data'].split(',')[1] if ',' in file['data'] else file['data']
                        image_bytes = base64.b64decode(base64_data)
                        st.image(image_bytes, caption=file['name'], width=200)
                    except:
                        st.write(f"📎 {file['name']}")
            
            # Display voice input information
            if message.get("audio_metadata") and message["audio_metadata"]["used_voice_input"]:
                st.caption(f"🎤 Voice input ({message['audio_metadata']['transcription_method']})")


# Clear history
if st.button("Clear History"):
    st.session_state.chat_history = []
    st.rerun()

Example Chat App

For a complete example implementation, visit the GitHub repository.

License

MIT License

Author

Jon Goncalves - Quadis

Links

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

quadis_chat_input-1.0.0.tar.gz (133.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

quadis_chat_input-1.0.0-py3-none-any.whl (132.2 kB view details)

Uploaded Python 3

File details

Details for the file quadis_chat_input-1.0.0.tar.gz.

File metadata

  • Download URL: quadis_chat_input-1.0.0.tar.gz
  • Upload date:
  • Size: 133.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.6

File hashes

Hashes for quadis_chat_input-1.0.0.tar.gz
Algorithm Hash digest
SHA256 8e2bf768881d319b934fb665e259c98aebad562ca4ad2678771852ec0e6f7688
MD5 44e7ed936c4cac0c2d3d0d61b39b6eee
BLAKE2b-256 4d22e55718818d451021a65944c6fb823cd308b39cfbe942ff33d6168594d733

See more details on using hashes here.

File details

Details for the file quadis_chat_input-1.0.0-py3-none-any.whl.

File metadata

File hashes

Hashes for quadis_chat_input-1.0.0-py3-none-any.whl
Algorithm Hash digest
SHA256 2fa2f147e5055578726ecb67c58e63bcf125ba65ed8b17cf9328242f14dd813b
MD5 5ad4b2e32b573b9ed5fe5c2125a603de
BLAKE2b-256 04ae0a066c713c7d6e7d0cbf664557f69c3a43852742897ae3f14444d08588ff

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page