<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://irislab.cau.ac.kr/feed.xml" rel="self" type="application/atom+xml" /><link href="https://irislab.cau.ac.kr/" rel="alternate" type="text/html" /><updated>2026-07-22T02:27:31+00:00</updated><id>https://irislab.cau.ac.kr/feed.xml</id><title type="html">IRIS Lab @ Chung-Ang University</title><subtitle>Immersive Reality &amp; Integrated Systems Lab</subtitle><author><name>Hak Gu Kim</name></author><entry><title type="html">Our paper accepted to ICIP 2026 Workshop</title><link href="https://irislab.cau.ac.kr/2026/06/27/workshop-icipw26.html" rel="alternate" type="text/html" title="Our paper accepted to ICIP 2026 Workshop" /><published>2026-06-27T00:00:00+00:00</published><updated>2026-06-27T00:00:00+00:00</updated><id>https://irislab.cau.ac.kr/2026/06/27/workshop-icipw26</id><content type="html" xml:base="https://irislab.cau.ac.kr/2026/06/27/workshop-icipw26.html"><![CDATA[<p>Congratualtions!</p>

<p>Our paper has been accepted to the <em>Many Lenses, One World: Culturally-Aware Interactive AI in the Metaverse Workshop</em> at the <strong>IEEE International Conference on Image Processing (ICIP) 2026</strong> <a href="https://2026.ieeeicip.org/">[LINK]</a></p>

<ul>
  <li>
    <p><strong>Title</strong>: Revealing Hidden Response Ambiguity in Binary Evaluation of Cultural Gesture Understanding</p>
  </li>
  <li>
    <p><strong>Authors</strong>: Sunghun Kang, Seungjae Lee, Hongseok Cho, Hyeokjun Kweon, Hak Gu Kim</p>
  </li>
  <li>
    <p><strong>Abstract</strong>: Gestures are culturally dependent forms of non-verbal communication whose meanings can vary substantially across regions. Recent benchmarks such as MC-SIGNS evaluate whether vision-language models (VLMs) correctly judge the appropriateness of gestures across cultural contexts, often by parsing responses into binary Yes/No outcomes. However, such binary evaluation can obscure important differences in the underlying response forms. In this paper, we investigate the hidden ambiguity underlying parsed “No” responses in culturally sensitive gesture understanding. We show that identical binary outcomes can arise from fundamentally different response forms, including culturally grounded rejection, over-refusal, and self-contradictory reasoning. To analyze this, we introduce a post-hoc response categorization framework that decomposes parsed “No” responses into interpretable categories using lexical cues in raw outputs. Experiments on three open-weight VLMs show that binary evaluation alone can substantially overestimate culturally informed gesture understanding by conflating valid negative judgments with spurious response forms.</p>
  </li>
</ul>]]></content><author><name>Hak Gu Kim</name></author><summary type="html"><![CDATA[Congratualtions!]]></summary></entry><entry><title type="html">Our paper accepted to Interspeech 2026 (Top-tier Conference)</title><link href="https://irislab.cau.ac.kr/2026/06/04/conf-interspeech26.html" rel="alternate" type="text/html" title="Our paper accepted to Interspeech 2026 (Top-tier Conference)" /><published>2026-06-04T00:00:00+00:00</published><updated>2026-06-04T00:00:00+00:00</updated><id>https://irislab.cau.ac.kr/2026/06/04/conf-interspeech26</id><content type="html" xml:base="https://irislab.cau.ac.kr/2026/06/04/conf-interspeech26.html"><![CDATA[<p>Congratualtions!</p>

<p>Our paper has been accepted to the <strong>Interspeech 2026</strong> <em>(Top-tier Conference)</em> <a href="https://interspeech2026.org/en-AU">[LINK]</a></p>

<ul>
  <li>
    <p><strong>Title</strong>: ArtBoost: Synthetic Articulatory Data Augmentation for Acoustic-to-Articulatory Inversion</p>
  </li>
  <li>
    <p><strong>Authors</strong>: Hyung Kyu Kim, Byungchan Hwang, and Hak Gu Kim</p>
  </li>
  <li>
    <p><strong>Abstract</strong>: Recent acoustic-to-articulatory inversion (AAI) models rely on electromagnetic articulography (EMA) data, which are costly and limited in scale. To address this limitation, we propose ArtBoost, a novel data augmentation strategy that leverages large-scale speech–mesh datasets originally developed for speech-driven 3D facial animation to improve AAI under limited EMA supervision. ArtBoost extracts pseudo articulatory trajectories from visible facial anchors and uses them for pre-training before fine-tuning on real EMA data. Experiments show consistent improvements in PCC and RMSE. Trajectory analyses confirm that the pseudo articulatory signals reflect physically meaningful visible articulatory dynamics. Additional evaluations across different AAI architectures demonstrate stable performance gains, indicating that ArtBoost can be integrated into diverse AAI models. These results suggest that speech–mesh data provide an effective and scalable source of articulatory supervision for AAI.</p>
  </li>
</ul>]]></content><author><name>Hak Gu Kim</name></author><summary type="html"><![CDATA[Congratualtions!]]></summary></entry><entry><title type="html">Our paper accepted to IEEE Trans. Information Forensics and Security (TIFS) (JCR Top 7.8%)</title><link href="https://irislab.cau.ac.kr/2026/05/19/journal-ieee-tifs.html" rel="alternate" type="text/html" title="Our paper accepted to IEEE Trans. Information Forensics and Security (TIFS) (JCR Top 7.8%)" /><published>2026-05-19T00:00:00+00:00</published><updated>2026-05-19T00:00:00+00:00</updated><id>https://irislab.cau.ac.kr/2026/05/19/journal-ieee%20tifs</id><content type="html" xml:base="https://irislab.cau.ac.kr/2026/05/19/journal-ieee-tifs.html"><![CDATA[<p>Congratualtions!</p>

<p>Our paper has been accepted to the <strong>IEEE Trans. Information Forensics and Security (TIFS)</strong> <em>(JCR Top 7.8%, Impact Factor: 8.0)</em> <a href="https://ieeexplore.ieee.org/document/11534487">[LINK]</a></p>

<ul>
  <li>
    <p><strong>Title</strong>: Deformable 3D Point Cloud Perturbations using Cage-based Deformation for Semantic Consistency</p>
  </li>
  <li>
    <p><strong>Authors</strong>: Kyo Seok Lee and Hak Gu Kim</p>
  </li>
  <li>
    <p><strong>Abstract</strong>: Deep neural networks for 3D point cloud analysis are widely used in applications such as autonomous driving and robotics, yet they remain highly vulnerable to adversarial attacks. Existing methods typically minimize point-wise distances to preserve geometry, which constrains perturbations and leads to a trade-off between imperceptibility and attack strength. To address this limitation, we propose a cage-based adversarial deformation framework that generates semantically consistent perturbations aligned with natural intra-class variations. Our method refines a source cage, predicts adversarial cage displace ments by fusing source–target features, and computes smooth point-wise offsets using solid-angle– and distance-aware weights. This enables globally coherent deformations that appear natural to humans while effectively misleading classifiers. Experiments on ModelNet40 and ShapeNet-Part show that our approach achieves state-of-the-art attack success rates while producing the most uniform point distributions and lowest local distortions. Furthermore, the perturbations remain effective against common defenses such as SRS, SOR, and DUP-Net.</p>
  </li>
</ul>]]></content><author><name>Hak Gu Kim</name></author><summary type="html"><![CDATA[Congratualtions!]]></summary></entry><entry><title type="html">Our paper accepted to IEEE Signal Processing Letters</title><link href="https://irislab.cau.ac.kr/2025/12/27/journal-ieee-spl.html" rel="alternate" type="text/html" title="Our paper accepted to IEEE Signal Processing Letters" /><published>2025-12-27T00:00:00+00:00</published><updated>2025-12-27T00:00:00+00:00</updated><id>https://irislab.cau.ac.kr/2025/12/27/journal-ieee-spl</id><content type="html" xml:base="https://irislab.cau.ac.kr/2025/12/27/journal-ieee-spl.html"><![CDATA[<p>Congratualtions!</p>

<p>Our paper has been accepted to the <strong>IEEE Signal Processing Letters (SPL)</strong> (Impact Factor: 3.9) <a href="https://ieeexplore.ieee.org/document/11320277">[LINK]</a></p>

<ul>
  <li>
    <p><strong>Title</strong>: Unsharp-Inspired Adversarial Point Cloud Perturbation via Low-Rank Approximation</p>
  </li>
  <li>
    <p><strong>Authors</strong>: Kyo Seok Lee, Han-nyoung Lee, and Hak Gu Kim</p>
  </li>
  <li>
    <p><strong>Abstract</strong>: Deep neural networks (DNNs) for 3D point cloud recognition have achieved remarkable performance, but remain highly vulnerable to adversarial perturbations. Existing adversarial methods often suffer from a trade-off between imperceptibility and attack strength, either producing noticeable outliers or requiring excessive point displacements. In this paper, we propose a novel unsharp-inspired adversarial perturbation that leverages low-rank approximation to balance global structure preservation and fine-detail manipulation. Specifically, the point cloud is decomposed via eigenvalue decomposition (EVD) into global and residual components, where the low-rank reconstruction captures the overall shape and the residual highlights salient details. Perturbations are then selectively applied to these fine-detail regions and smoothed according to local magnitude and orientation, ensuring geometric consistency. Extensive experiments on ModelNet40 and ShapeNet-Part using PointNet and DGCNN demonstrate that our method achieves 100\% attack success rate while significantly reducing geometric distortion compared to state-of-the-art methods. These results highlight the effectiveness of spectral-domain low-rank modeling for generating adversarial point clouds that are both strong and imperceptible.</p>
  </li>
</ul>]]></content><author><name>Hak Gu Kim</name></author><summary type="html"><![CDATA[Congratualtions!]]></summary></entry><entry><title type="html">Seungjae Lee received the Best Student Paper Award at the IEIE Summer Conference 2025</title><link href="https://irislab.cau.ac.kr/2025/06/27/award-ieie-summer26.html" rel="alternate" type="text/html" title="Seungjae Lee received the Best Student Paper Award at the IEIE Summer Conference 2025" /><published>2025-06-27T00:00:00+00:00</published><updated>2025-06-27T00:00:00+00:00</updated><id>https://irislab.cau.ac.kr/2025/06/27/award-ieie-summer26</id><content type="html" xml:base="https://irislab.cau.ac.kr/2025/06/27/award-ieie-summer26.html"><![CDATA[<p>Congratualtions!</p>

<p>Our Master’s student Seungjae Lee has won the <strong>Best Student Paper Award at IEIE Summer Conference 2025</strong> <a href="http://conf2025s.ieieweb.org/2025s/">[LINK]</a></p>

<ul>
  <li>
    <p><strong>Title</strong>: 3D Head and Body Pose Estimation 기반 갑상선 환자 재활운동 정량적 평가 및 안내 시스템</p>
  </li>
  <li>
    <p><strong>Authors</strong>: 이승재, 김호준, 이하린, 김태용, 김학구</p>
  </li>
  <li>
    <p><strong>Abstract</strong>: With the rising incidence of thyroid cancer in Korea, postoperative neck rehabilitation has become increasingly important. However, many patients struggle to follow physician-recommended exercises due to pain and lack of real-time monitoring, leading to delayed recovery. To address this problem, in this paper, we propose a deep learning-based automatic rehabilitation exercise guide system that quantitatively evaluates rehabilitation performance of a patient using 3D head and body pose estimation with real-time feedback. The proposed system also integrates a metaverse-based environment, enabling patients to perform guided exercises remotely without visiting the hospital.</p>
  </li>
</ul>

<div class="news-image">
  <img src="/assets/images/news/250627_ieie_summer25.png" alt="IEIE Summer Conference 2025" />
</div>]]></content><author><name>Hak Gu Kim</name></author><summary type="html"><![CDATA[Congratualtions!]]></summary></entry><entry><title type="html">Our paper accepted to ICCV 2025 (Top-tier Conference)</title><link href="https://irislab.cau.ac.kr/2025/06/26/conf-iccv25.html" rel="alternate" type="text/html" title="Our paper accepted to ICCV 2025 (Top-tier Conference)" /><published>2025-06-26T00:00:00+00:00</published><updated>2025-06-26T00:00:00+00:00</updated><id>https://irislab.cau.ac.kr/2025/06/26/conf-iccv25</id><content type="html" xml:base="https://irislab.cau.ac.kr/2025/06/26/conf-iccv25.html"><![CDATA[<p>Congratualtions!</p>

<p>Our paper has been accepted to the <strong>IEEE/CVF International Conference on Computer Vision (ICCV) 2025</strong> <em>(Top-tier Conference)</em> <a href="https://iccv.thecvf.com/">[LINK]</a></p>

<ul>
  <li>
    <p><strong>Title</strong>: MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided Stylization</p>
  </li>
  <li>
    <p><strong>Authors</strong>: Hyung Kyu Kim, Sangmin Lee (Korea University), and Hak Gu Kim</p>
  </li>
  <li>
    <p><strong>Abstract</strong>: Speech-driven 3D facial animation aims to synthesize realistic facial motion sequences from given audio, matching the speaker’s speaking style. However, previous works often require priors such as class labels of a speaker or additional 3D facial meshes at inference, which makes them fail to reflect the speaking style and limits their practical use. To address these issues, we propose MemoryTalker which enables realistic and accurate 3D facial motion synthesis by reflecting speaking style only with audio input to maximize usability in applications. Our framework consists of two training stages: &lt;1-stage&gt; is storing and retrieving general motion (i.e., Memorizing), and &lt;2-stage&gt; is to perform the personalized facial motion synthesis (i.e., Animating) with the motion memory stylized by the audio-driven speaking style feature. In this second stage, our model learns about which facial motion types should be emphasized for a particular piece of audio. As a result, our MemoryTalker can generate a reliable personalized facial animation without additional prior information.  With quantitative and qualitative evaluations, as well as user study, we show the effectiveness of our model and its performance enhancement for personalized facial animation over state-of-the-art methods. Our source code will be released to facilitate further research.</p>
  </li>
</ul>]]></content><author><name>Hak Gu Kim</name></author><summary type="html"><![CDATA[Congratualtions!]]></summary></entry><entry><title type="html">Our paper accepted to ICIP 2025</title><link href="https://irislab.cau.ac.kr/2025/05/21/conf-icip25.html" rel="alternate" type="text/html" title="Our paper accepted to ICIP 2025" /><published>2025-05-21T00:00:00+00:00</published><updated>2025-05-21T00:00:00+00:00</updated><id>https://irislab.cau.ac.kr/2025/05/21/conf-icip25</id><content type="html" xml:base="https://irislab.cau.ac.kr/2025/05/21/conf-icip25.html"><![CDATA[<p>Congratualtions!</p>

<p>Our paper has been accepted to the <strong>IEEE International Conference on Image Processing (ICIP) 2025</strong> <a href="https://2025.ieeeicip.org/">[LINK]</a></p>

<ul>
  <li>
    <p><strong>Title</strong>: Enhancing 3D Scene Representation with Structural Dissimilarity-Aware Learning</p>
  </li>
  <li>
    <p><strong>Authors</strong>: Seungjae Lee, Ho Jun Kim, and Hak Gu Kim</p>
  </li>
  <li>
    <p><strong>Abstract</strong>: Novel view synthesis aims to generate high-quality unseen views from images at different viewpoints. However, the existing methods often struggle to preserve fine details, leading to structural distortions in complex regions. In this paper, we introduce a simple yet effective structure-aware objective function designed to enhance structural information in novel view synthesis. By leveraging the Structural Similarity Index (SSIM), our method attends to regions exhibiting significant structural distortions. We incorporate structural dissimilarity-based attention to highlight discrepancies in challenging regions between predicted and ground-truth images. It enables recent 3D scene representation models to achieve improved structure preservation, leading to more coherent representations. Experiments on synthetic and real-world datasets demonstrate that our approach enhances structural consistency, particularly in challenging regions, making it a valuable addition to state-of-the-arts.</p>
  </li>
</ul>]]></content><author><name>Hak Gu Kim</name></author><summary type="html"><![CDATA[Congratualtions!]]></summary></entry><entry><title type="html">Our paper accepted to Interspeech 2025 (Top-tier Conference)</title><link href="https://irislab.cau.ac.kr/2025/05/19/conf-interspeech25.html" rel="alternate" type="text/html" title="Our paper accepted to Interspeech 2025 (Top-tier Conference)" /><published>2025-05-19T00:00:00+00:00</published><updated>2025-05-19T00:00:00+00:00</updated><id>https://irislab.cau.ac.kr/2025/05/19/conf-interspeech25</id><content type="html" xml:base="https://irislab.cau.ac.kr/2025/05/19/conf-interspeech25.html"><![CDATA[<p>Congratualtions!</p>

<p>Our paper has been accepted to the <strong>Interspeech 2025</strong> <em>(Top-tier Conference)</em> <a href="https://interspeech2025.org/en-AU">[LINK]</a></p>

<ul>
  <li>
    <p><strong>Title</strong>: Learning Phonetic Context-Dependent Viseme for Enhancing Speech-Driven 3D Facial Animation</p>
  </li>
  <li>
    <p><strong>Authors</strong>: Hyung Kyu Kim and Hak Gu Kim</p>
  </li>
  <li>
    <p><strong>Abstract</strong>: Speech-driven 3D facial animation aims to generate realistic facial movements synchronized with audio. Traditional methods primarily minimize reconstruction loss by aligning each frame with ground-truth. However, this frame-wise approach often fails to capture the continuity of facial motion, leading to jittery and unnatural outputs due to coarticulation. To address this, we propose a novel phonetic context-aware loss, which explicitly models the influence of phonetic context on viseme transitions. By incorporating a viseme coarticulation weight, we assigns adaptive importance to facial movements based on their dynamic changes over time, ensuring smoother and perceptually consistent animations. Extensive experiments demonstrate that replacing the conventional reconstruction loss with ours improves both quantitative metrics and visual quality. It highlights the importance of explicitly modeling phonetic context dependent visemes in synthesizing natural speech-driven 3D facial animation.</p>
  </li>
</ul>]]></content><author><name>Hak Gu Kim</name></author><summary type="html"><![CDATA[Congratualtions!]]></summary></entry><entry><title type="html">Our paper accepted to IEEE Access</title><link href="https://irislab.cau.ac.kr/2025/04/29/journal-ieee-access.html" rel="alternate" type="text/html" title="Our paper accepted to IEEE Access" /><published>2025-04-29T00:00:00+00:00</published><updated>2025-04-29T00:00:00+00:00</updated><id>https://irislab.cau.ac.kr/2025/04/29/journal-ieee-access</id><content type="html" xml:base="https://irislab.cau.ac.kr/2025/04/29/journal-ieee-access.html"><![CDATA[<p>Congratualtions!</p>

<p>Our paper has been accepted to the <strong>IEEE Access</strong> (Impact Factor: 4.2) <a href="https://ieeexplore.ieee.org/document/10981835">[LINK]</a></p>

<ul>
  <li>
    <p><strong>Title</strong>: Leveraging Text Signed Distance Function Map for Boundary-Aware Guidance in Scene Text Segmentation</p>
  </li>
  <li>
    <p><strong>Authors</strong>: Ho Jun Kim and Hak Gu Kim</p>
  </li>
  <li>
    <p><strong>Abstract</strong>: Scene text segmentation is to predict pixel-wise text regions from an image, enabling in-image text editing or removal. One of the primary challenges is to remove noises including non-text regions and predict intricate text boundaries. To deal with that, traditional approaches utilize a text detection or recognition module explicitly. However, they are likely to highlight noise around the text. Because they did not sufficiently consider the boundaries of text, they fail to accurately predict the fine details of text. In this paper, we introduce leveraging text signed distance function (SDF) map, which encodes distance information from text boundaries, in scene text segmentation to explicitly provide text boundary information. By spatial cross attention mechanism, we encode the text-attended feature from the text SDF map. Then, both visual and text-attended features are utilized to decode the text segmentation map. Our approach not only mitigates confusion between text and complex backgrounds by eliminating false positives such as logos and texture blobs located far from the text, but also effectively captures fine details of complex text patterns by leveraging text boundary information. Extensive experiments demonstrate that leveraging text SDF map in scene text segmentation provides superior performances on various scene text segmentation datasets.</p>
  </li>
</ul>]]></content><author><name>Hak Gu Kim</name></author><summary type="html"><![CDATA[Congratualtions!]]></summary></entry><entry><title type="html">Our workshop proposal accepted to ICIP 2025</title><link href="https://irislab.cau.ac.kr/2025/02/15/workshop-ieee-icip25.html" rel="alternate" type="text/html" title="Our workshop proposal accepted to ICIP 2025" /><published>2025-02-15T00:00:00+00:00</published><updated>2025-02-15T00:00:00+00:00</updated><id>https://irislab.cau.ac.kr/2025/02/15/workshop-ieee-icip25</id><content type="html" xml:base="https://irislab.cau.ac.kr/2025/02/15/workshop-ieee-icip25.html"><![CDATA[<p>Congratualtions!</p>

<p>Our workshop proposal has been accepted to the <strong>IEEE International Conference on Image Processing (ICIP) 2025</strong> <a href="https://carai.kaist.ac.kr/lvlm">[LINK]</a></p>

<ul>
  <li>
    <p><strong>Title</strong>: 2nd Integrating Image Processing with Large-Scale Language/Vision Models for Advanced Visual Understanding</p>
  </li>
  <li>
    <p><strong>Authors</strong>: Yong Mak Ro (KAIST), Wen-Huang Cheng (National Taiwan Univ.), and Hak Gu Kim (Chung-Ang Univ.)</p>
  </li>
  <li>
    <p><strong>Abstract</strong>: This workshop aims to bridge the gap between conventional image processing techniques and the latest advancements in large-scale vision and language models. Recent developments in large-scale models have revolutionized image processing tasks, significantly enhancing capabilities in visual object understanding, image classification, and generative image synthesis. Furthermore, the large-scale models have opened new avenues for human-machine multimodal interactive dialogue systems, where the synergy between visual and linguistic processing enables more intuitive and dynamic interactions. This workshop will provide a platform for researchers and practitioners to explore how cutting-edge large-scale models integrate with image processing methods and foster innovation across diverse applications. Discussions will extend beyond conventional tasks to address the role of vision-language models in Generative AI and their use in multimodal systems, such as virtual assistants that interact seamlessly using images, text, and speech.</p>
  </li>
</ul>]]></content><author><name>Hak Gu Kim</name></author><summary type="html"><![CDATA[Congratualtions!]]></summary></entry></feed>