Abstract
Math Information Retrieval (MIR) aims to revolutionize the retrieval of math information by addressing the challenges in representing, retrieving, and producing highly specialized math content. In order to develop a robust MIR system, designers must be able to process math equations (MEs) to a format that different systems can use, which is difficult due to various formats math information are stored in, including visual images and document texts. To improve the usage of math information in varied forms for people who are learning and utilizing math, we propose three different methods of handling ME data. The first method is a ME extraction system that (i) applies a one-shot object detector to identify math equations in digital images using an efficient neural architecture search method and (ii) employs a sequence-to-sequence (Seq2Seq) encoder-decoder system to recognize math equation symbols based on the Bayesian Neural Network (BNN) row encoding. This system balances speed and accuracy of a Math IR system for the purpose of extracting math information from images so that any text processor system can handle math content in visual mediums. The second method is using Document to Vector (Doc2Vec) and Formula to Vector (Formula2Vec) to determine content and formula similarity among math questions and their potential answers, respectively. The proposed model, two Vector Information Retrieval (2VecIR), applies both Doc2Vec and Formula2Vec in tandem, which enables a deep understanding of both the language within the text and the numeric or symbolic content of the formulas and determines the most relevant information to return to the user. We have also designed a Math Large Language Model (LLM) based on NExT-GPT with a custom core architecture based off of Llemma.
Degree
MS
College and Department
Computational, Mathematical, and Physical Sciences; Computer Science
Rights
https://lib.byu.edu/about/copyright/
BYU ScholarsArchive Citation
Crawford, Angel Wheelwright, "Math Information Retrieval from Images via Formula Extraction, Vector Embedding, and Multimodal LLM-Based Question Answering" (2025). Theses and Dissertations. 11375.
https://scholarsarchive.byu.edu/etd/11375
Date Submitted
2025-07-23
Document Type
Thesis
Keywords
image detection, document and formula vectors, LLMs, math question answering
Language
english