Mothman at SemEval-2024 Task 9: An Iterative System for Chain-of-Thought Prompt Optimization

Chen, Alvin Po-Chun, Groshan, Ray, von Bayern, Sean

May-3-2024–arXiv.org Artificial Intelligence

Extensive research exists on the performance of large language models on logic-based tasks, whereas relatively little has been done on their ability to generate creative solutions on lateral thinking tasks. The BrainTeaser shared task tests lateral thinking and uses adversarial datasets to prevent memorization, resulting in poor performance for out-of-the-box models. We propose a system for iterative, chain-of-thought prompt engineering which optimizes prompts using human evaluation. Using this shared task, we demonstrate our system's ability to significantly improve model performance by optimizing prompts and evaluate the input dataset.

bridge, dataset, hay, (14 more...)

arXiv.org Artificial Intelligence

May-3-2024

arXiv.org PDF

Add feedback

Country:
- Europe > Italy (0.04)
- Asia > Singapore (0.04)
- North America
  - United States
    - New York (0.05)
    - Colorado > Boulder County
      - Boulder (0.04)
  - Mexico > Mexico City
    - Mexico City (0.04)

Genre:
- Research Report > Promising Solution (0.48)

Technology:
- Information Technology > Artificial Intelligence > Natural Language > Large Language Model (1.00)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found