← 論文一覧に戻る

GreenBench: Apple Silicon上でのオープンソースLLM推論のエネルギー効率とカーボンフットプリントのベンチマーキング

GreenBench: Benchmarking Energy Efficiency and Carbon Footprint of Open-Source LLM Inference on Apple Silicon (原題)

Rajeswari Kannan, Raj Firke, Shreya Bengle, Srushti Deshmukh

arXiv (Cornell University)プレプリント2026-08-24#AI×ESG経営インパクト: コスト削減対象セクター: cross_sector
DOI: 10.48550/arxiv.2608.28667
原典: https://arxiv.org/abs/2608.28667
📄 PDF

🤖 gxceed AI 要約

日本語

本論文は、Apple Silicon(M4 Pro)上でのオープンソースLLM推論のエネルギー効率と炭素排出を評価するベンチマークフレームワークGreenBenchを提案する。実測により、単一ユーザー利用ではデータセンターGPU比で30〜40倍のエネルギー効率を達成し、モデルサイズと効率のトレードオフを明らかにした。

English

This paper presents GreenBench, a benchmarking framework to evaluate energy efficiency and carbon footprint of open-source LLM inference on Apple Silicon. Measurements show 30-40x better energy efficiency per token than datacenter GPUs for single-user deployment, and identify optimal model choices for accuracy-efficiency trade-offs.

Unofficial AI-generated summary based on the public title and abstract. Not an official translation.

📝 gxceed 編集解説 — Why this matters

日本のGX文脈において

日本では、AI活用拡大に伴うエネルギー消費と脱炭素の両立が課題。本研究成果は、エッジAIやオンプレミスLLM導入時のエネルギー効率評価に資する。

In the global GX context

Globally, this contributes to Green AI by extending energy efficiency benchmarking beyond datacenters to edge devices, informing sustainable AI deployment strategies and carbon footprint accounting.

👥 読者別の含意

🔬研究者:Provides a methodology and baseline for energy-efficient LLM inference on edge hardware.

🏢実務担当者:Offers guidance for selecting models and hardware to reduce energy costs and carbon footprint in AI deployments.

📄 Abstract(原文)

The rapid proliferation of Large Language Models (LLMs) has raised concerns about their environmental impact during inference. While Green AI research has focused on datacenter GPUs and embedded platforms, the energy profile of LLM inference on Apple Silicon, with its unified memory architecture, remains unstudied. This paper presents GreenBench, a benchmarking framework that evaluates the energy efficiency, throughput, and carbon footprint of five open-source LLMs (3-9B parameters) across three NLP tasks on an Apple M4 Pro with 48 GB unified memory. Using macOS powermetrics for direct power measurement and Ollama's nanosecond-precision timing, we find that the M4 Pro draws only 0.47 W of CPU+GPU package power during sustained inference, with total system power of 8-12 W, achieving 30-40x better energy efficiency per token than datacenter GPUs in single-user deployment. Smaller models (3-3.8B) deliver 2.6-4.2x higher throughput and up to 62% less energy per token than larger models (7-9B). Pareto analysis identifies Qwen 2.5 (7B) as the optimal accuracy-efficiency trade-off at 57% MMLU and 59 tokens/s, while Llama 3.2 (3B) suits latency-critical applications at 175 tokens/s. We provide per-token energy at package and system levels with CO2 estimates for India and US grids.

🔗 Provenance — このレコードを発見したソース

🔔 こうした論文の新着を逃したくない方は キーワードアラート に登録(無料・3キーワードまで)。

gxceed は公開メタデータに基づく研究支援データセットです。要約・翻訳・解説は AI 支援で生成されています。 最終的な解釈・検証は利用者が原典資料に基づいて行うことを前提とします。