Raven / README.md
Ruvadev's picture
Update README.md
26c3baa verified
|
Raw History Blame Contribute Delete
4.99 kB
metadata
library_name: pytorch
pipeline_tag: image-classification
tags:
  - computer-vision
  - image-forensics
  - ai-image-detection
  - dinov2
  - safetensors

Raven

Raven์€ ์‹ค์ œ ์‚ฌ์ง„๊ณผ AI ์ƒ์„ฑ ์ด๋ฏธ์ง€๋ฅผ ๊ตฌ๋ถ„ํ•˜๊ธฐ ์œ„ํ•ด ๋งŒ๋“  ์ด๋ฏธ์ง€ ํฌ๋ Œ์‹ ๋ชจ๋ธ์ž…๋‹ˆ๋‹ค.

์ตœ์ข… ์ถœ๋ ฅ์€ REAL, AI, UNCERTAIN ๋กœ ์ด 3๊ฐ€์ง€์ด๊ณ , AI ๋ฐ์ดํ„ฐ๋Š” GPT Image 2 ๊ณ„์—ด์„ ๊ธฐ์ค€์œผ๋กœ ํ•ฉ๋‹ˆ๋‹ค.

ํŒ์ • ๊ธฐ์ค€์€ ๋‹ค์Œ๊ณผ ๊ฐ™์Šต๋‹ˆ๋‹ค.

REAL       p(AI) <= 0.28
UNCERTAIN  0.28 < p(AI) < 0.72
AI         p(AI) >= 0.72

์˜ˆ์‹œ์ฝ”๋“œ

import warnings
warnings.filterwarnings("ignore")

import sys
from pathlib import Path

HERE = Path(__file__).resolve().parent
sys.path.insert(0, str(HERE))

from raven.inference import infer_image

model = HERE / "model.safetensors"

image_extensions = {
    ".png",
    ".jpg",
    ".jpeg",
    ".webp",
    ".bmp",
}

images = sorted(
    [
        file for file in HERE.iterdir()
        if file.is_file()
        and file.suffix.lower() in image_extensions
    ],
    key=lambda p: p.name.lower()
)

for image in images:
    print(f"File: {image.name}")

    result = infer_image(
        str(model),
        str(image),
    )

    print(f"Verdict: {result['verdict']}")
    print(f"AI: {result['ai_probability'] * 100:.2f}%")
    print(f"REAL: {result['real_probability'] * 100:.2f}%")
    print()

Benchmark

Validation ๋ฐ์ดํ„ฐ๋Š” ์ด 4,470์žฅ์ž…๋‹ˆ๋‹ค.

  • REAL: 3,003
  • AI: 1,467
Metric Result 95% CI
Accuracy 98.635% 98.251% - 98.936%
Balanced Accuracy 98.566% 98.159% - 98.935%
AUROC 0.998255 0.997220 - 0.999083
Balanced AP 0.998472 0.997685 - 0.999135
AI confirmed recall 97.001% 95.998% - 97.758%
REAL confirmed recall 97.502% 96.881% - 98.003%
AI to REAL error 0.954% 0.569% - 1.596%
REAL to AI error 0.599% 0.379% - 0.946%
Coverage 98.054% -
Selective accuracy 99.207% -
Uncertain 1.946% -
Balanced Brier 0.011403 -
Balanced ECE 0.005572 -

Decisions

AI 1,467์žฅ:

AI             1,423
REAL              14
UNCERTAIN         30

REAL 3,003์žฅ:

REAL           2,928
AI                18
UNCERTAIN         57

REAL-only Test

๋ณ„๋„๋กœ ๋ถ„๋ฆฌ๋œ REAL ์ด๋ฏธ์ง€ 2,986์žฅ์—์„œ๋„ ํ‰๊ฐ€๋ฅผ ํ•˜์˜€์Šต๋‹ˆ๋‹ค.

Metric Result 95% CI
Accuracy 98.225% 97.686% - 98.640%
REAL confirmed recall 96.383% 95.652% - 96.995%
REAL to AI error 1.038% 0.732% - 1.470%
Coverage 97.421% -
Selective accuracy 98.934% -
Uncertain 2.579% -

์‹ค์ œ ํŒ์ • ๊ฒฐ๊ณผ:

REAL           2,878
AI                31
UNCERTAIN         77

Lighting

Validation ๋ฐ์ดํ„ฐ์˜ ๋ฐ๊ธฐ๋ณ„ ๊ฒฐ๊ณผ์ž…๋‹ˆ๋‹ค.

Group N Accuracy AUROC AI Recall REAL Recall Uncertain
dark 299 99.331% 0.999498 98.611% 92.771% 2.676%
dim 1,137 98.769% 0.999256 97.767% 97.139% 2.199%
extreme-dark 28 96.429% 1.000000 94.444% 90.000% 7.143%
normal 3,006 98.536% 0.997519 96.265% 97.840% 1.730%

extreme-dark๋Š” ํ‘œ๋ณธ ์ˆ˜๊ฐ€ ์ ๊ธฐ ๋•Œ๋ฌธ์— ๋‹ค๋ฅธ ๊ตฌ๊ฐ„๋ณด๋‹ค ์ˆ˜์น˜์˜ ๋ถˆํ™•์‹ค์„ฑ์ด ํด ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Limitations

ํ˜„์žฌ AI ํ•™์Šต ๋ฐ์ดํ„ฐ๋Š” GPT Image 2๋ฅผ ์ค‘์‹ฌ์œผ๋กœ ๊ตฌ์„ฑ๋˜์–ด ์žˆ์Šต๋‹ˆ๋‹ค.

๋”ฐ๋ผ์„œ ์œ„ benchmark๋Š” GPT Image 2 ๊ณ„์—ด๊ณผ ํ˜„์žฌ REAL ๋ฐ์ดํ„ฐ ๋ถ„ํฌ์—์„œ์˜ ์„ฑ๋Šฅ์„ ๋‚˜ํƒ€๋ƒ…๋‹ˆ๋‹ค.

๋‹ค์Œ๊ณผ ๊ฐ™์€ ๊ฒฝ์šฐ ๋™์ผํ•œ ์„ฑ๋Šฅ์„ ๋ณด์žฅํ•˜์ง€ ์•Š์Šต๋‹ˆ๋‹ค.

  • ํ•™์Šต์— ํฌํ•จ๋˜์ง€ ์•Š์€ AI ์ƒ์„ฑ ๋ชจ๋ธ
  • ๊ฐ•ํ•œ JPEG ์žฌ์••์ถ•
  • ์Šคํฌ๋ฆฐ์ƒท
  • ์—…์Šค์ผ€์ผ ๋ฐ ๋…ธ์ด์ฆˆ ์ œ๊ฑฐ
  • ๊ณผ๋„ํ•œ ์ƒ‰๋ณด์ • ๋˜๋Š” ํ›„์ฒ˜๋ฆฌ
  • ์ด๋ฏธ์ง€ ์ผ๋ถ€๋งŒ ํ•ฉ์„ฑ๋œ ๊ฒฝ์šฐ
  • ๋งค์šฐ ์–ด๋‘์šด ์ด๋ฏธ์ง€

ํŠนํžˆ Midjourney, FLUX, Stable Diffusion ๋“ฑ ๋‹ค๋ฅธ ์ƒ์„ฑ๊ธฐ์—์„œ์˜ ์„ฑ๋Šฅ์€ ๋ณ„๋„๋กœ ๊ฒ€์ฆ๋˜์–ด์•ผ ํ•ฉ๋‹ˆ๋‹ค.

Raven์˜ ์ถœ๋ ฅ์€ ์ด๋ฏธ์ง€ ์ถœ์ฒ˜์— ๋Œ€ํ•œ ํ™•๋ฅ  ๊ธฐ๋ฐ˜ ํฌ๋ Œ์‹ ํŒ๋‹จ์ด๋ฉฐ, ์ด๋ฏธ์ง€๊ฐ€ AI๋กœ ์ƒ์„ฑ๋˜์—ˆ์Œ์„ ์ฆ๋ช…ํ•˜๋Š” ์ ˆ๋Œ€์ ์ธ ์ฆ๊ฑฐ๋กœ ์‚ฌ์šฉํ•ด์„œ๋Š” ์•ˆ ๋ฉ๋‹ˆ๋‹ค.