Picture for Minchan Kim

Minchan Kim

Spatio-Temporal Graphs Beyond Grids: Benchmark for Maritime Anomaly Detection

Add code
Dec 23, 2025
Figure 1 for Spatio-Temporal Graphs Beyond Grids: Benchmark for Maritime Anomaly Detection
Figure 2 for Spatio-Temporal Graphs Beyond Grids: Benchmark for Maritime Anomaly Detection
Viaarxiv icon

Adaptive Sparsified Graph Learning Framework for Vessel Behavior Anomalies

Add code
Feb 20, 2025
Viaarxiv icon

SegINR: Segment-wise Implicit Neural Representation for Sequence Alignment in Neural Text-to-Speech

Add code
Oct 07, 2024
Figure 1 for SegINR: Segment-wise Implicit Neural Representation for Sequence Alignment in Neural Text-to-Speech
Figure 2 for SegINR: Segment-wise Implicit Neural Representation for Sequence Alignment in Neural Text-to-Speech
Figure 3 for SegINR: Segment-wise Implicit Neural Representation for Sequence Alignment in Neural Text-to-Speech
Figure 4 for SegINR: Segment-wise Implicit Neural Representation for Sequence Alignment in Neural Text-to-Speech
Viaarxiv icon

CANVAS: Commonsense-Aware Navigation System for Intuitive Human-Robot Interaction

Add code
Oct 02, 2024
Figure 1 for CANVAS: Commonsense-Aware Navigation System for Intuitive Human-Robot Interaction
Figure 2 for CANVAS: Commonsense-Aware Navigation System for Intuitive Human-Robot Interaction
Figure 3 for CANVAS: Commonsense-Aware Navigation System for Intuitive Human-Robot Interaction
Figure 4 for CANVAS: Commonsense-Aware Navigation System for Intuitive Human-Robot Interaction
Viaarxiv icon

High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model

Add code
Jun 25, 2024
Figure 1 for High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model
Figure 2 for High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model
Figure 3 for High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model
Figure 4 for High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model
Viaarxiv icon

MakeSinger: A Semi-Supervised Training Method for Data-Efficient Singing Voice Synthesis via Classifier-free Diffusion Guidance

Add code
Jun 10, 2024
Figure 1 for MakeSinger: A Semi-Supervised Training Method for Data-Efficient Singing Voice Synthesis via Classifier-free Diffusion Guidance
Figure 2 for MakeSinger: A Semi-Supervised Training Method for Data-Efficient Singing Voice Synthesis via Classifier-free Diffusion Guidance
Figure 3 for MakeSinger: A Semi-Supervised Training Method for Data-Efficient Singing Voice Synthesis via Classifier-free Diffusion Guidance
Viaarxiv icon

Exploiting Semantic Reconstruction to Mitigate Hallucinations in Vision-Language Models

Add code
Mar 26, 2024
Viaarxiv icon

Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction

Add code
Jan 03, 2024
Figure 1 for Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction
Figure 2 for Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction
Figure 3 for Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction
Figure 4 for Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction
Viaarxiv icon

Efficient Parallel Audio Generation using Group Masked Language Modeling

Add code
Jan 02, 2024
Figure 1 for Efficient Parallel Audio Generation using Group Masked Language Modeling
Figure 2 for Efficient Parallel Audio Generation using Group Masked Language Modeling
Figure 3 for Efficient Parallel Audio Generation using Group Masked Language Modeling
Figure 4 for Efficient Parallel Audio Generation using Group Masked Language Modeling
Viaarxiv icon

Transduce and Speak: Neural Transducer for Text-to-Speech with Semantic Token Prediction

Add code
Nov 08, 2023
Viaarxiv icon