Skip to main content

Creating Unit Tests using OpenAI

Introduction

The original intent of this codebase was to perform prompt engineering via "vectorization" of a java codebase and then feeding the embedded text to openAI for it to automatically generate unit tests. More languages and LLMs will eventually be supported, and the use cases aren't necessarily limited to unit test generation.

This repository contains several unrelated/experimental files based on past iterations, but in general the module lives in the src/llm_prompt_creator directory.

The instructions in this README are kept up to date as much as possible.

Contributing

Note that the main branch is locked down but does allow merge requests.

To contribute, create a feature or fix branch (prepended with feature_ or fix_ respectively), commit your changes there and then create a pull request from your branch into main.

We will review & (after approval) merge your git branch and then delete the remote branch on our github repo to limit left-over branches.

Set Up

Note Windows users may need to install the Visual Studio C++ Compiler to use this package

Simple example usage:

# Below is an example on how to set the OpenAI key,
# it has to be above the "langchain" and "llm_prompt_creator" import.
# Create an "openai-key.txt" in the same directory as your test.py file.

import os

with open('openai-key.txt', 'r') as f:
    key = f.read().strip()
    os.environ["OPENAI_API_KEY"] = key


from langchain.chat_models import ChatOpenAI
from llm_prompt_creator import prompt as PR

from llm_prompt_creator import prompt as PR
dir = "<path to your java codebase directory>"

# Chunk & store your codebase as tokenized chunks via javalang.
# Defaults to store succesfully chunked files in "./chunks.json".

PR.chunker(dir)

"""
You could optionally store the chunks strictly in memory by instead using the below when chunking your
directory:
"""
#data = PR.chunker(dir, write_to_disk=False)

"""
Create a vector store to perform a similarity search against when asking questions to your
LLM. Defaults to consume from the "./chunks.json" file.
"""
store = PR.create_vectorstore()

"""
If opting to save the store to disk, use the below instead which passes a
directory where the store will be saved. It will also load the store into
memory for follow on commands.
"""
#PR.create_vectorstore(persist_directory="db")
#store = PR.load_vectorstore(persist_directory="db")

# Start an open-ended chat conversation with your LLM based on your vector store.
# Will continue prompting the user for inputs until they type 'exit'.
# Subject to model limitations (especially token limits).
PR.prompt(store=store, llm=ChatOpenAI(model="gpt-4",temperature=0))

"""
To show the context provided (provided by the vector store based on the user's question)
uncomment the below:
"""
#PR.prompt(store, show_context=True)

"""
To not write the accumulated context to disk while still displaying context in terminal, use the below:
"""

#PR.prompt(store, show_context=True, write_to_disk=False)

"""
To provide a custom prompt template or a list of questions to be automatically prompted for, use the filePath parameter.
The file should be a json file with properties of promptTemplate and questions. An example file can be found below:
"""
{
"promptTemplate": "",
"questions": ["question 1", "question 2"]
}


#PR.prompt(store, show_context=True, filePath="./file_input.json", llm=ChatOpenAI(model="gpt-4",temperature=0))

Following the example should yield a similar response to the below image (subject to LLM model used and codebase):

TODO

  • Refactoring across the board, particularly to reduce the number of called Python scripts.
  • Optimize chunker to allow larger codebase directories
  • Establish a standard way of calculating token limits.
  • Use token limits to dynamically adjust the amount of context and therefore the number of tokens used during a prompt/completion instance with OpenAI.
  • Containerize this solution so we can deploy it; one for parsing and chunking, another for creating a vector-store and prompting (or something like it).

Metadata

Release files for llm-prompt-creator 0.5.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llm-prompt-creator 0.5.0
File Size Uploaded
llm_prompt_creator-0.5.0.tar.gz 7.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llm-prompt-creator 0.5.0
File Interpreter ABI Platform
llm_prompt_creator-0.5.0-py3-none-any.whl Python 3 none any Details

Total release size: 15.4 kB

Release files / llm_prompt_creator-0.5.0.tar.gz

Download URL llm_prompt_creator-0.5.0.tar.gz
Size 7.5 kB
Tags Source
SHA-256 checksum
How to use checksums
68431a1d71744e531d69387504a592d7d91f41c4d0834f960c358c21867704eb
BLAKE2b-256 checksum
How to use checksums
0cca2ae2823c71715578190ad7d6cd6c3e7be8a7812d00b9c2b172d784711e1e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.11.4

Release files / llm_prompt_creator-0.5.0-py3-none-any.whl

Download URL llm_prompt_creator-0.5.0-py3-none-any.whl
Size 7.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c064a34c7143acf078273acc670210b3e1b5ed07830d25426f0b8256f913ae44
BLAKE2b-256 checksum
How to use checksums
0c4ff9d60e54bde3e06783a7eba3c91a0b9949ca14826e45627f52043cdf3d7a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.11.4

Release history Release notifications | RSS feed

This release

0.5.0 This release

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.16

2 release files

0.2.15

2 release files

0.2.14

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page