Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Xue-Yong Fu

Beyond Text-to-SQL: An Agentic LLM System for Governed Enterprise Analytics APIs

May 20, 2026

Gundeep Singh, Parsa Kavehzadeh, Jing Xia, Xue-Yong Fu, Julien Bouvier Tremblay, Md Tahmid Rahman Laskar, Vincent Lum, Shashi Bhushan TN

Abstract:Enterprise analytics aims to make organizational data accessible for decision-making, yet non-technical users still face barriers when using traditional business intelligence tools or Text-to-SQL systems. While recent Text-to-SQL approaches based on Large Language Models (LLMs) promise natural language access to structured data, they fall short in enterprise settings where analytics pipelines rely on governed APIs rather than raw databases. In practice, these APIs encapsulate complex business logic to ensure consistency, auditability, and security. However, delegating mathematical or aggregation logic to an LLM introduces reliability and compliance risks. To this end, we present Analytic Agent, an LLM-based agentic system that translates natural language intents into secure interactions with enterprise analytics APIs. Evaluated on 90 real enterprise use cases constructed by domain experts, it reliably interprets user goals, validates permissions, executes governed queries, and generates compliant visualizations through multi-step reasoning and policy-aware orchestration.

* The first four authors contributed equally to this work

Via

Access Paper or Ask Questions

From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents

May 14, 2026

Md Tahmid Rahman Laskar, Xue-Yong Fu, Seyyed Saeed Sarfjoo, Quinten McNamara, Jonas Robertson, Shashi Bhushan TN

Abstract:Voice agents increasingly require reliable tool use from speech, whereas prominent tool-calling benchmarks remain text-based. We study whether verified text benchmarks can be converted into controlled audio-based tool calling evaluations without re-annotating the tool schema and gold labels. Our dataset-agnostic framework uses text-to-speech, speaker variation, and environmental noise to create paired text-audio instances while preserving the original dataset annotations. Based on extensive evaluation of 7 omni-modal models on audio-converted versions of Confetti and When2Call, our framework demonstrates that the performance is strongly model- and task-dependent: Gemini-3.1-Flash-Live obtains the highest Confetti score (70.4), whereas GPT-Realtime-1.5 performs best on When2Call (71.9). On Confetti, the text-to-voice gap ranges from 1.8 points for Qwen3-Omni to 4.8 points for GPT-Realtime-1.5. A targeted analysis of failure cases demonstrates that degradations most often reflect misunderstandings of argument values in the speech. Considering real-world deployment scenarios, we further report text-only results, an ambiguity-based reformulation stress test, and a reference-free LLM-as-judge protocol validated against human preferences. Notably, we find that open-source Qwen3 judges with at least 8B parameters exceed 80% agreement with proprietary judges, supporting privacy-preserving evaluation. Overall, our framework provides a verifiable and reproducible first-stage diagnostic that complements purpose-built audio corpora.

Via

Access Paper or Ask Questions

Query-OPT: Optimizing Inference of Large Language Models via Multi-Query Instructions in Meeting Summarization

Feb 29, 2024

Md Tahmid Rahman Laskar, Elena Khasanova, Xue-Yong Fu, Cheng Chen, Shashi Bhushan TN

Figure 1 for Query-OPT: Optimizing Inference of Large Language Models via Multi-Query Instructions in Meeting Summarization

Figure 2 for Query-OPT: Optimizing Inference of Large Language Models via Multi-Query Instructions in Meeting Summarization

Figure 3 for Query-OPT: Optimizing Inference of Large Language Models via Multi-Query Instructions in Meeting Summarization

Figure 4 for Query-OPT: Optimizing Inference of Large Language Models via Multi-Query Instructions in Meeting Summarization

Abstract:This work focuses on the task of query-based meeting summarization in which the summary of a context (meeting transcript) is generated in response to a specific query. When using Large Language Models (LLMs) for this task, a new call to the LLM inference endpoint/API is required for each new query even if the context stays the same. However, repeated calls to the LLM inference endpoints would significantly increase the costs of using them in production, making LLMs impractical for many real-world use cases. To address this problem, in this paper, we investigate whether combining the queries for the same input context in a single prompt to minimize repeated calls can be successfully used in meeting summarization. In this regard, we conduct extensive experiments by comparing the performance of various popular LLMs: GPT-4, PaLM-2, LLaMA-2, Mistral, and FLAN-T5 in single-query and multi-query settings. We observe that while most LLMs tend to respond to the multi-query instructions, almost all of them (except GPT-4), even after fine-tuning, could not properly generate the response in the required output format. We conclude that while multi-query prompting could be useful to optimize the inference costs by reducing calls to the inference endpoints/APIs for the task of meeting summarization, this capability to reliably generate the response in the expected format is only limited to certain LLMs.

Via

Access Paper or Ask Questions

Tiny Titans: Can Smaller Large Language Models Punch Above Their Weight in the Real World for Meeting Summarization?

Feb 01, 2024

Xue-Yong Fu, Md Tahmid Rahman Laskar, Elena Khasanova, Cheng Chen, Shashi Bhushan TN

Figure 1 for Tiny Titans: Can Smaller Large Language Models Punch Above Their Weight in the Real World for Meeting Summarization?

Figure 2 for Tiny Titans: Can Smaller Large Language Models Punch Above Their Weight in the Real World for Meeting Summarization?

Figure 3 for Tiny Titans: Can Smaller Large Language Models Punch Above Their Weight in the Real World for Meeting Summarization?

Figure 4 for Tiny Titans: Can Smaller Large Language Models Punch Above Their Weight in the Real World for Meeting Summarization?

Abstract:Large Language Models (LLMs) have demonstrated impressive capabilities to solve a wide range of tasks without being explicitly fine-tuned on task-specific datasets. However, deploying LLMs in the real world is not trivial, as it requires substantial computing resources. In this paper, we investigate whether smaller, compact LLMs are a good alternative to the comparatively Larger LLMs2 to address significant costs associated with utilizing LLMs in the real world. In this regard, we study the meeting summarization task in a real-world industrial environment and conduct extensive experiments by comparing the performance of fine-tuned compact LLMs (e.g., FLAN-T5, TinyLLaMA, LiteLLaMA) with zero-shot larger LLMs (e.g., LLaMA-2, GPT-3.5, PaLM-2). We observe that most smaller LLMs, even after fine-tuning, fail to outperform larger zero-shot LLMs in meeting summarization datasets. However, a notable exception is FLAN-T5 (780M parameters), which performs on par or even better than many zero-shot Larger LLMs (from 7B to above 70B parameters), while being significantly smaller. This makes compact LLMs like FLAN-T5 a suitable cost-efficient solution for real-world industrial deployment.

* The first two authors contributed equally to this work

Via

Access Paper or Ask Questions

Building Real-World Meeting Summarization Systems using Large Language Models: A Practical Perspective

Nov 08, 2023

Md Tahmid Rahman Laskar, Xue-Yong Fu, Cheng Chen, Shashi Bhushan TN

Figure 1 for Building Real-World Meeting Summarization Systems using Large Language Models: A Practical Perspective

Figure 2 for Building Real-World Meeting Summarization Systems using Large Language Models: A Practical Perspective

Figure 3 for Building Real-World Meeting Summarization Systems using Large Language Models: A Practical Perspective

Figure 4 for Building Real-World Meeting Summarization Systems using Large Language Models: A Practical Perspective

Abstract:This paper studies how to effectively build meeting summarization systems for real-world usage using large language models (LLMs). For this purpose, we conduct an extensive evaluation and comparison of various closed-source and open-source LLMs, namely, GPT-4, GPT- 3.5, PaLM-2, and LLaMA-2. Our findings reveal that most closed-source LLMs are generally better in terms of performance. However, much smaller open-source models like LLaMA- 2 (7B and 13B) could still achieve performance comparable to the large closed-source models even in zero-shot scenarios. Considering the privacy concerns of closed-source models for only being accessible via API, alongside the high cost associated with using fine-tuned versions of the closed-source models, the opensource models that can achieve competitive performance are more advantageous for industrial use. Balancing performance with associated costs and privacy concerns, the LLaMA-2-7B model looks more promising for industrial usage. In sum, this paper offers practical insights on using LLMs for real-world business meeting summarization, shedding light on the trade-offs between performance and cost.

* EMNLP 2023 Industry Track

Via

Access Paper or Ask Questions

Are Large Language Models Reliable Judges? A Study on the Factuality Evaluation Capabilities of LLMs

Nov 01, 2023

Xue-Yong Fu, Md Tahmid Rahman Laskar, Cheng Chen, Shashi Bhushan TN

Figure 1 for Are Large Language Models Reliable Judges? A Study on the Factuality Evaluation Capabilities of LLMs

Figure 2 for Are Large Language Models Reliable Judges? A Study on the Factuality Evaluation Capabilities of LLMs

Figure 3 for Are Large Language Models Reliable Judges? A Study on the Factuality Evaluation Capabilities of LLMs

Abstract:In recent years, Large Language Models (LLMs) have gained immense attention due to their notable emergent capabilities, surpassing those seen in earlier language models. A particularly intriguing application of LLMs is their role as evaluators for texts produced by various generative models. In this study, we delve into the potential of LLMs as reliable assessors of factual consistency in summaries generated by text-generation models. Initially, we introduce an innovative approach for factuality assessment using LLMs. This entails employing a singular LLM for the entirety of the question-answering-based factuality scoring process. Following this, we examine the efficacy of various LLMs in direct factuality scoring, benchmarking them against traditional measures and human annotations. Contrary to initial expectations, our results indicate a lack of significant correlations between factuality metrics and human evaluations, specifically for GPT-4 and PaLM-2. Notable correlations were only observed with GPT-3.5 across two factuality subcategories. These consistent findings across various factual error categories suggest a fundamental limitation in the current LLMs' capability to accurately gauge factuality. This version presents the information more concisely while maintaining the main points and findings of the original text.

* accepted by Generation, Evaluation & Metrics (GEM) Workshop at EMNLP 2023

Via

Access Paper or Ask Questions

AI Coach Assist: An Automated Approach for Call Recommendation in Contact Centers for Agent Coaching

May 28, 2023

Md Tahmid Rahman Laskar, Cheng Chen, Xue-Yong Fu, Mahsa Azizi, Shashi Bhushan, Simon Corston-Oliver

Figure 1 for AI Coach Assist: An Automated Approach for Call Recommendation in Contact Centers for Agent Coaching

Figure 2 for AI Coach Assist: An Automated Approach for Call Recommendation in Contact Centers for Agent Coaching

Figure 3 for AI Coach Assist: An Automated Approach for Call Recommendation in Contact Centers for Agent Coaching

Figure 4 for AI Coach Assist: An Automated Approach for Call Recommendation in Contact Centers for Agent Coaching

Abstract:In recent years, the utilization of Artificial Intelligence (AI) in the contact center industry is on the rise. One area where AI can have a significant impact is in the coaching of contact center agents. By analyzing call transcripts using Natural Language Processing (NLP) techniques, it would be possible to quickly determine which calls are most relevant for coaching purposes. In this paper, we present AI Coach Assist, which leverages the pre-trained transformer-based language models to determine whether a given call is coachable or not based on the quality assurance (QA) questions asked by the contact center managers or supervisors. The system was trained and evaluated on a large dataset collected from real-world contact centers and provides an effective way to recommend calls to the contact center managers that are more likely to contain coachable moments. Our experimental findings demonstrate the potential of AI Coach Assist to improve the coaching process, resulting in enhancing the performance of contact center agents.

* ACL 2023 Industry Track

Via

Access Paper or Ask Questions

Improving Named Entity Recognition in Telephone Conversations via Effective Active Learning with Human in the Loop

Nov 02, 2022

Md Tahmid Rahman Laskar, Cheng Chen, Xue-Yong Fu, Shashi Bhushan TN

Figure 1 for Improving Named Entity Recognition in Telephone Conversations via Effective Active Learning with Human in the Loop

Figure 2 for Improving Named Entity Recognition in Telephone Conversations via Effective Active Learning with Human in the Loop

Figure 3 for Improving Named Entity Recognition in Telephone Conversations via Effective Active Learning with Human in the Loop

Figure 4 for Improving Named Entity Recognition in Telephone Conversations via Effective Active Learning with Human in the Loop

Abstract:Telephone transcription data can be very noisy due to speech recognition errors, disfluencies, etc. Not only that annotating such data is very challenging for the annotators, but also such data may have lots of annotation errors even after the annotation job is completed, resulting in a very poor model performance. In this paper, we present an active learning framework that leverages human in the loop learning to identify data samples from the annotated dataset for re-annotation that are more likely to contain annotation errors. In this way, we largely reduce the need for data re-annotation for the whole dataset. We conduct extensive experiments with our proposed approach for Named Entity Recognition and observe that by re-annotating only about 6% training instances out of the whole dataset, the F1 score for a certain entity type can be significantly improved by about 25%.

* The final version of this paper will be published in the Proceedings of the DaSH Workshop @ EMNLP 2022. This paper is accepted for presentation in both DaSH@EMNLP 2022 and HiLL@NIPS 2022

Via

Access Paper or Ask Questions

Entity-level Sentiment Analysis in Contact Center Telephone Conversations

Oct 26, 2022

Xue-Yong Fu, Cheng Chen, Md Tahmid Rahman Laskar, Shayna Gardiner, Pooja Hiranandani, Shashi Bhushan TN

Figure 1 for Entity-level Sentiment Analysis in Contact Center Telephone Conversations

Figure 2 for Entity-level Sentiment Analysis in Contact Center Telephone Conversations

Figure 3 for Entity-level Sentiment Analysis in Contact Center Telephone Conversations

Figure 4 for Entity-level Sentiment Analysis in Contact Center Telephone Conversations

Abstract:Entity-level sentiment analysis predicts the sentiment about entities mentioned in a given text. It is very useful in a business context to understand user emotions towards certain entities, such as products or companies. In this paper, we demonstrate how we developed an entity-level sentiment analysis system that analyzes English telephone conversation transcripts in contact centers to provide business insight. We present two approaches, one entirely based on the transformer-based DistilBERT model, and another that uses a convolutional neural network supplemented with some heuristic rules.

* EMNLP 2022

Via

Access Paper or Ask Questions

An Effective, Performant Named Entity Recognition System for Noisy Business Telephone Conversation Transcripts

Sep 27, 2022

Xue-Yong Fu, Cheng Chen, Md Tahmid Rahman Laskar, Shashi Bhushan TN, Simon Corston-Oliver

Figure 1 for An Effective, Performant Named Entity Recognition System for Noisy Business Telephone Conversation Transcripts

Figure 2 for An Effective, Performant Named Entity Recognition System for Noisy Business Telephone Conversation Transcripts

Figure 3 for An Effective, Performant Named Entity Recognition System for Noisy Business Telephone Conversation Transcripts

Figure 4 for An Effective, Performant Named Entity Recognition System for Noisy Business Telephone Conversation Transcripts

Abstract:We present a simple yet effective method to train a named entity recognition (NER) model that operates on business telephone conversation transcripts that contain noise due to the nature of spoken conversation and artifacts of automatic speech recognition. We first fine-tune LUKE, a state-of-the-art Named Entity Recognition (NER) model, on a limited amount of transcripts, then use it as the teacher model to teach a smaller DistilBERT-based student model using a large amount of weakly labeled data and a small amount of human-annotated data. The model achieves high accuracy while also satisfying the practical constraints for inclusion in a commercial telephony product: realtime performance when deployed on cost-effective CPUs rather than GPUs.

Via

Access Paper or Ask Questions