ラベル Hardware Performance Counter の投稿を表示しています。 すべての投稿を表示
ラベル Hardware Performance Counter の投稿を表示しています。 すべての投稿を表示

2008/12/14

Intel CPU: Fixed Performance Counter (FPC)

Fixed Performance Counter(FPC) // 勝手な命名
は固定のパフォーマンスイベントをカウントするMSR.
最大で3つで,CPUがいくつサポートしているかは,以前のエントリ
パフォーマンスカウンタの情報を表示
の "Number of fixed-function PCs" で得られる.

以下FPCの詳細.

[IA32_FIXED_CTR0(MSR_PERF_FIXED_CTR0)]
Register Address: 0x309
Performance Event: Inst_Retired.Any
実行した命令数をカウントする.

[IA32_FIXED_CTR1(MSR_PERF_FIXED_CTR1)]
Register Address: 0x30A
Performance Event: CPU_CLK_Unhalted.Core
CPUがHalt状態でない時のCPUサイクル数をカウントする.

[IA32_FIXED_CTR2(MSR_PERF_FIXED_CTR2)]
Register Address: 0x30B
Performance Event: CPU_CLK_Unhalted.Ref
CPUがHalt状態でない時のバスクロック数をカウントする.


【使用時の注意】
FPCはデフォルトではdisableになっている.
enableにするには,FPCの動作をコントロールするMSRである
IA32_FIXED_CTR_CTL(MSR_PERF_FIXED_CTR_CTL)
Register Address: 0x3BD
の値を変更する必要がある.
IA32_FIXED_CTR_CTLは,3つのFPCのenable/disableの設定を行う.


[bit fields]
0-4: IA32_FIXED_CTR0の設定
0: EN0_OS(CPL=0の時にenable)
1: EN0_Usr(CPL>0の時にenable)
2: Reserved
3: EN0_PMI(カウンタがオーバーフローした時にPMIを発生)
4-7: IA32_FIXED_CTR1の設定,0-3と同様
8-11: IA32_FIXED_CTR2の設定,0-3と同様


[Reference]
IntelR 64 and IA-32 Architectures Software Developer’s Manual Volume 3B: APPENDIX B
MODEL-SPECIFIC REGISTERS (MSRS)

2008/12/11

Intel CPU: パフォーマンスイベントUmask

モニタリングするパフォーマンスイベントのUmaskでTable 18-11, 18-13, 18-14を見るべし.
という指示がちょくちょくあったのでメモ.
intel manual Vol. 3 capter 18 p59-60 より,


[Table 18-11. Core Specificity Encoding]
Bit 15:14
11B All cores
10B Reserved
01B This core
00B Reserved

Some microarchitectural conditions allow detection specificity only at the boundary of physical processors. Some bus events belong to this category, providing specificity between the originating physical processor (a bus agent) versus other agents on the bus. Sub-field encoding for agent specificity is shown in Table 18-12.


[Table 18-12. Agent Specificity Encoding]
Bit 13
0 This agent
1 Include all agents

Some microarchitectural conditions are detectable only from the originating core. In such cases, unit mask does not support core-specificity or agent-specificity encodings. These are referred to as core-only conditions. Some microarchitectural conditions allow detection specificity that includes or excludes the action of hardware prefetches. A two-bit encoding may be supported to qualify hardware prefetch actions. Typically, this applies only to some L2 or bus events. The sub-field encoding for hardware prefetch qualification is shown in Table 18-13.

[Table 18-13. HW Prefetch Qualification Encoding]
Bit 13:12
11B All inclusive
10B Reserved
01B Hardware prefetch only
00B Exclude hardware prefetch

[Table 18-14. MESI Qualification Definitions]
Bit Position 11:8
Bit 11 Counts modified state
Bit 10 Counts exclusive state
Bit 9 Counts shared state
Bit 8 Counts Invalid state

// よく使いそうな16進マスク
//--------------------------------------------------
[Table 18-11. Core Specificity Encoding]
0xc000: All cores
0x4000: This core

[Table 18-12. Agent Specificity Encoding]
0x0000 This agent
0x2000 Include all agents

[Table 18-13. HW Prefetch Qualification Encoding]
0x3000: All inclusive
0x1000: Hardware prefetch only

[Table 18-14. MESI Qualification Definitions]
0x0f00: 全部


// Reference
Intel® 64 and IA-32 Architectures Software Developer’s Manual Volume 3B: APPENDIX A
PERFORMANCE-MONITORING EVENTS

2008/12/10

C言語: ハードウェアパフォーマンスカウンタの情報を表示

CPUIDでハードウェアパフォーマンスカウンタの情報を表示するプログラム.
#include 
#include 

#define BIT(x, bit) (((x) >> (bit)) & 0x00000001)

int main()
{
  unsigned int a, b, c, d;
  unsigned char eax_07_00, eax_15_08, eax_23_16, eax_31_24;
  unsigned char edx_04_00, edx_12_05;

  asm volatile ("cpuid" : "=a"(a), "=b"(b), "=c"(c), "=d"(d) : "a"(0x0a));

  printf("CPUID(EAX=0AH): Architectural Performance Monitoring\n");

  eax_07_00 = (unsigned char)(a & 0xff);
  eax_15_08 = (unsigned char)((a >> 8) & 0xff);
  eax_23_16 = (unsigned char)((a >> 16) & 0xff);
  eax_31_24 = (unsigned char)((a >> 24) & 0xff);
  printf(" Version ID of architectural PM: %d\n", eax_07_00);
  printf(" Number of general-purpose PMC per logical processor: %d\n", eax_15_08);
  printf(" Bit width of general-purpose PMC: %d\n", eax_23_16);
  printf(" Length of EBX bit vector: %d\n", eax_31_24);
  printf("\n");
 
  printf(" Core cycle event not available: %d\n", BIT(b, 0));
  printf(" Instruction retired event not available: %d\n", BIT(b, 1));
  printf(" Reference cycles event not available: %d\n", BIT(b, 2));
  printf(" Last-level cache reference event not available: %d\n", BIT(b, 3));
  printf(" Last-level cache misses event not available: %d\n", BIT(b, 4));
  printf(" Branch instrunction retired event not available: %d\n", BIT(b, 5));
  printf(" Branch mispredict retired event not available: %d\n", BIT(b, 6));
  printf("\n");

  if (eax_07_00 > 1) {
    edx_04_00 = (unsigned char)(d & 0x1f);
    edx_12_05 = (unsigned char)((d >> 5) & 0xff);
    printf(" Number of fixed-function PCs: %d\n", edx_04_00);
    printf(" Bit width of fixed-function PCs: %d\n", edx_12_05);
  }

  return 0;
}

2008/12/01

Intel CPU: Hardware Performance Counter(HPC)

intel CPUのハードウェアパフォーマンスカウンタ(HPC)は
自分で監視するイベントを指定できる汎用的なカウンタと
固定のイベントを監視する固定カウンタの二種類がある.

Core 2では汎用カウンタは2個,固定カウンタは3個ある.
固定カウンタで監視するイベントは
Inst_Retired.Any: 完了した命令の数
CPU_CLK_Unhalted.Core: HLT状態でないCPUのクロックサイクル数
CPU_CLK_Unhalted.Ref: 上記のreference cycle版

【用語】
CCCR: Counter Configuration ContRol
ESCRで選択されたイベントをフィルタリングしてカウントする場合に使用する.
CPL: Current Privilege Level
ESCR: Event Selection ContRol
LLC: Last Level Cache
MSR: Model Specific Register
PEBS: Precise Event-Based Sampling
PMC: PerforMance Counter
PMI: Performance Monitoring Interrupt

【命令】
(RD/WR)MSR: MSRをRead/Writeする命令
RDPMC: PMCをReadする命令

【制限】
MSRをRead/Writeする命令を実行できるのは特権モード(ring 0)のみ.
アプリケーション(ring 3)からはPMCのReadのみ可能.

【イベント】
[リタイアイベント]
リタイアしたイベントとは,マシンステートを確定した命令によるイベントのみを含む.

[非リタイアイベント]
非リタイアイベントは,マイクロアーキテクチャーにおける
アウトオブオーダーの推測による流れで発生する.

[L2キャッシュイベント]
・MEM_LOAD_RETIRED.L2_LINE_MISS
L2キャッシュのミスで発生したロード数

・MEM_LOAD_RETIRED.L2_MISS


【参考資料】
インテルCoreマイクロアーキテクチャー・プロセッサーのパフォーマンス・カウンター
MEM_LOAD_RETIRED.L2_LINE_MISS
MEM_LOAD_RETIRED.L2_MISS

rdmsr - 通信用語の基礎知識
CPU-CPUID-MSR - SyncHack
CPU-MSR - SyncHack
linux2.6-include-asm-i386-msr.h - LinuxKernelHackJapan

Intel® 64 and IA-32 Architectures Software Developer's Manuals

2008/11/25

Xenoprof: Xenのパフォーマンスカウンタ

OProfile はカーネルも含めたシステム全体のプロファイリングを行うもの.
XenoprofはOProfileの拡張.

xenoprofを使うためには圧縮されていない,XenとLinuxのカーネルイメージが必要
(/boot/xen-syms-[version]とXenのコンパイルツリーのトップにあるvmlinux)

vmlinuxとvmlinuzは別物.
vmlinuzはvmlinuxを圧縮したもの?
展開の仕方は不明.

// OProfileを有効にする
# make menuconfig


Instrumentation Support --->
[*] Profiling support (EXPERIMENTAL)
OProfile system profiling (EXPERIMENTAL)

上記がチェックされていることを確認し,再コンパイル.

[opcontrol]
Active domains

--passive-domains
データを取るドメインUのIDを指定する.


[opreport]
opreport event:[EVENT]
指定したイベントのプロファイルのみ表示する.

opreport cpu:[number]
指定したCPUのプロファイルのみ表示する.


// Reference
[OProfile]
OProfileとは
Benchmark-OProfile - PukiWiki
Omicron OProfile
OProfileを使ってCPUプロファイリングをとる

OProfile documentation
OProfile manual

[xenoprof]
Opensource.hp.com - Welcome

[xenperf]
Linux Virtualization Guides - Xen 3.0 User Guide - Xen Build Options
Xen/Xen Tools/xenperf - PukiWiki

http://prdownloads.sourceforge.net/oprofile/oprofile-0.9.4.tar.gz