연구

연구

최정상 AI 학회에 15편 이상 논문을 발표하고 6건의 특허를 출원했습니다. 선도적인 학계 및 산업 파트너와 함께 연구하며, 모든 성과를 제품에 반영합니다.
논문
arXiv 2026*

EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards

1,200 robot-view video scenarios testing whether VLM guards catch real hazards without over-alarming

arXiv 2026*

When Context Flips, Safety Breaks: Diagnosing Brittle Safety in Aligned Language Models

Context-flip evaluation revealing models cling to safety rules even when the safe action becomes harmful

arXiv 2026*

XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity

Country-grounded safety benchmark with 5,500 cases across 10 country-language pairs

arXiv 2026*

PHASOR: Phase-Anchored Universal Action Representations for Humanoid Embodiments

Embodiment-agnostic action embeddings anchoring humanoid robots to a shared human motion manifold

KDD WS 2026

NRT-Bench: Benchmarking Multi-Turn Red-Teaming of LLM Operator Agents in Safety-Critical Control Rooms

Multi-turn adversarial red-teaming of LLM operator agents in a simulated nuclear control room

ECCV 2026

When Cars Have Stereotypes: Auditing Demographic Bias in Objects from Text-to-Image Models

SODA framework revealing demographic bias in generated objects across 8,000 images

WACV 2026

Better Safe Than Sorry? Overreaction Problem of Vision Language Models in Visual Emergency Recognition

VLMs misclassify 31–96% of safe situations as dangerous in visual emergency recognition

LREC 2026

Redefining Evaluation Standards: A Unified Framework for Evaluating the Korean Capabilities of Language Models

A unified framework for evaluating the Korean capabilities of language models

ICLR 2026

Jailbreaking on Text-to-Video Models via Scene Splitting Strategy

First systematic black-box jailbreak for T2V models, achieving 70–84% success rates

arXiv 2025*

When Good Sounds Go Adversarial: Jailbreaking Audio-Language Models with Benign Inputs

Jailbreaking audio-language models with benign-sounding adversarial inputs

ACL 2026

COMPASS: Evaluating Organization-Specific Policy Alignment in LLMs

Framework for evaluating LLM compliance with enterprise-specific allow/deny policies

ACL 2026

What Users Leave Unsaid: Under-Specified Queries Limit Vision-Language Models

HAERAE-Vision benchmark revealing VLM failures on ambiguous and incomplete queries

CVPR 2026

Fine-Grained Multi-Image Object Hallucination Benchmark

A comprehensive benchmark for evaluating object hallucination across multiple images in vision-language models

AISTATS 2026

Towards Motion-aware Referring Image Segmentation

Motion-aware approach for referring image segmentation with improved temporal understanding

Nature 2026

Humanity's Last Exam

Frontier AI evaluation benchmark with 3,000+ expert-level questions across disciplines

arXiv 2025

Eliciting and Analyzing Emergent Misalignment

Conversational red-teaming to elicit misalignment through narrative immersion and emotional pressure

ACL WS 2025

One-Shot is Enough: Consolidating Multi-Turn Attacks into Efficient Single-Turn Prompts for LLMs

Compressing multi-turn jailbreak strategies into single-turn attacks

ACL 2025

Representation Bending for Large Language Model Safety

Manipulating LLM internal representations to reduce harmful outputs

IEEE-EMBS 2025

PatientSafeBench: Evaluating Safety of Medical LLMs

Safety evaluation framework for patient-facing medical AI systems

ICML 2025

ELITE: Enhanced Language-Image Toxicity Evaluation

Benchmark for evaluating toxicity across language-image multimodal models

NeurIPS WS 2025

X-Teaming Evolutionary M2S

LLM-guided evolution for automated multi-turn jailbreak template discovery

NeurIPS WS 2025

ObjexMT: Objective Extraction & Metacognitive Calibration

Benchmarking whether LLM judges can recover hidden objectives in jailbreak transcripts

ACL 2025

sudo rm -rf agentic_security

Systematic security analysis of autonomous AI agent vulnerabilities · Industry Track

ICLR 2023

DepthFL: Depthwise Federated Learning

Federated learning for heterogeneous resource-constrained clients

* 심사 중 / 게재 목표 학회

특허
2025.12

AI Security Testing Integrated Framework and Visual Flow-Based Execution System

DB25-0017-KR0

2025.09

Multimodal Content Policy Violation Detection System

DB25-0016-KR0

2025.09

AI-Based NER Detection Guardrail

DB25-0015-KR0

2025.03

Domain-Specific AI Guardrail System

10-2025-0038904

2024.09

AI Guardrail System

10-2024-0124354

2024.08

Red Teaming Method for Security Assessment

10-2024-0116863

오픈소스
GitHub

Video2Robot

LinkedIn과 X에서 조회수 20만 회 이상, 스타 540개 이상

Hugging Face

AIM Intelligence

AI 안전성 연구를 위한 공개 모델 및 데이터셋

AI 안전성 연구를 함께 합시다.

공동 논문 발표부터 엔터프라이즈 보안 평가까지, AI를 더 안전하게 만드는 길을 함께 걷겠습니다.

전문가와 상담하기Hugging Face에서 보기
aim

AI를 지킬 준비가 되셨나요?

에임인텔리전스 보안 전문가와 상담하고, 시스템에 최적화된 무료 레드티밍 데모를 요청하세요.

플랫폼 살펴보기