2009-05-26 15:17:18 +08:00
|
|
|
perf-report(1)
|
2009-05-30 18:38:51 +08:00
|
|
|
==============
|
2009-05-26 15:17:18 +08:00
|
|
|
|
|
|
|
NAME
|
|
|
|
----
|
2009-05-27 15:33:18 +08:00
|
|
|
perf-report - Read perf.data (created by perf record) and display the profile
|
2009-05-26 15:17:18 +08:00
|
|
|
|
|
|
|
SYNOPSIS
|
|
|
|
--------
|
|
|
|
[verse]
|
|
|
|
'perf report' [-i <file> | --input=file]
|
|
|
|
|
|
|
|
DESCRIPTION
|
|
|
|
-----------
|
|
|
|
This command displays the performance counter profile information recorded
|
2009-06-23 22:39:53 +08:00
|
|
|
via perf record.
|
2009-05-26 15:17:18 +08:00
|
|
|
|
|
|
|
OPTIONS
|
|
|
|
-------
|
|
|
|
-i::
|
|
|
|
--input=::
|
2009-05-27 15:33:18 +08:00
|
|
|
Input file name. (default: perf.data)
|
2010-12-01 09:57:17 +08:00
|
|
|
|
|
|
|
-v::
|
|
|
|
--verbose::
|
|
|
|
Be more verbose. (show symbol address, etc)
|
|
|
|
|
2009-07-01 06:01:20 +08:00
|
|
|
-d::
|
|
|
|
--dsos=::
|
|
|
|
Only consider symbols in these dsos. CSV that understands
|
|
|
|
file://filename entries.
|
2009-11-09 19:26:13 +08:00
|
|
|
-n::
|
|
|
|
--show-nr-samples::
|
2009-07-11 23:18:37 +08:00
|
|
|
Show the number of samples for each symbol
|
2010-12-01 09:57:17 +08:00
|
|
|
|
|
|
|
--showcpuutilization::
|
|
|
|
Show sample percentage for different cpu modes.
|
|
|
|
|
2009-11-09 19:26:13 +08:00
|
|
|
-T::
|
|
|
|
--threads::
|
2009-08-07 19:55:24 +08:00
|
|
|
Show per-thread event counters
|
2011-11-14 02:30:08 +08:00
|
|
|
-c::
|
2009-07-01 06:01:21 +08:00
|
|
|
--comms=::
|
|
|
|
Only consider symbols in these comms. CSV that understands
|
|
|
|
file://filename entries.
|
2009-07-01 06:01:22 +08:00
|
|
|
-S::
|
|
|
|
--symbols=::
|
|
|
|
Only consider these symbols. CSV that understands
|
|
|
|
file://filename entries.
|
2009-05-26 15:17:18 +08:00
|
|
|
|
2010-12-01 09:57:17 +08:00
|
|
|
-U::
|
|
|
|
--hide-unresolved::
|
|
|
|
Only display entries resolved to a symbol.
|
|
|
|
|
perf diff: Use perf_session__fprintf_hists just like 'perf record'
That means that almost everything you can do with 'perf report'
can be done with 'perf diff', for instance:
$ perf record -f find / > /dev/null
[ perf record: Woken up 1 times to write data ]
[ perf record: Captured and wrote 0.062 MB perf.data (~2699
samples) ] $ perf record -f find / > /dev/null
[ perf record: Woken up 1 times to write data ]
[ perf record: Captured and wrote 0.062 MB perf.data (~2687
samples) ] perf diff | head -8
9.02% +1.00% find libc-2.10.1.so [.] _IO_vfprintf_internal
2.91% -1.00% find [kernel] [k] __kmalloc
2.85% -1.00% find [kernel] [k] ext4_htree_store_dirent
1.99% -1.00% find [kernel] [k] _atomic_dec_and_lock
2.44% find [kernel] [k] half_md4_transform
$
So if you want to zoom into libc:
$ perf diff --dsos libc-2.10.1.so | head -8
37.34% find [.] _IO_vfprintf_internal
10.34% find [.] __GI_memmove
8.25% +2.00% find [.] _int_malloc
5.07% -1.00% find [.] __GI_mempcpy
7.62% +2.00% find [.] _int_free
$
And if there were multiple commands using libc, it is also
possible to aggregate them all by using --sort symbol:
$ perf diff --dsos libc-2.10.1.so --sort symbol | head -8
37.34% [.] _IO_vfprintf_internal
10.34% [.] __GI_memmove
8.25% +2.00% [.] _int_malloc
5.07% -1.00% [.] __GI_mempcpy
7.62% +2.00% [.] _int_free
$
The displacement column now is off by default, to use it:
perf diff -m --dsos libc-2.10.1.so --sort symbol | head -8
37.34% [.] _IO_vfprintf_internal
10.34% [.] __GI_memmove
8.25% +2.00% [.] _int_malloc
5.07% -1.00% +2 [.] __GI_mempcpy
7.62% +2.00% -1 [.] _int_free
$
Using -t/--field-separator can be used for scripting:
$ perf diff -t, -m --dsos libc-2.10.1.so --sort symbol | head -8
37.34, , ,[.] _IO_vfprintf_internal
10.34, , ,[.] __GI_memmove
8.25,+2.00%, ,[.] _int_malloc
5.07,-1.00%, +2,[.] __GI_mempcpy
7.62,+2.00%, -1,[.] _int_free
6.99,+1.00%, -1,[.] _IO_new_file_xsputn
1.89,-2.00%, +4,[.] __readdir64
$
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
Cc: Frédéric Weisbecker <fweisbec@gmail.com>
Cc: Mike Galbraith <efault@gmx.de>
Cc: Peter Zijlstra <a.p.zijlstra@chello.nl>
Cc: Paul Mackerras <paulus@samba.org>
LKML-Reference: <1260978567-550-1-git-send-email-acme@infradead.org>
Signed-off-by: Ingo Molnar <mingo@elte.hu>
2009-12-16 23:49:27 +08:00
|
|
|
-s::
|
|
|
|
--sort=::
|
|
|
|
Sort by key(s): pid, comm, dso, symbol, parent.
|
|
|
|
|
2010-12-01 09:57:17 +08:00
|
|
|
-p::
|
|
|
|
--parent=<regex>::
|
|
|
|
regex filter to identify parent, see: '--sort parent'
|
|
|
|
|
|
|
|
-x::
|
|
|
|
--exclude-other::
|
|
|
|
Only display entries with parent-match.
|
|
|
|
|
2009-07-11 09:47:28 +08:00
|
|
|
-w::
|
2010-12-01 09:57:17 +08:00
|
|
|
--column-widths=<width[,width...]>::
|
2009-07-11 09:47:28 +08:00
|
|
|
Force each column width to the provided list, for large terminal
|
|
|
|
readability.
|
|
|
|
|
|
|
|
-t::
|
|
|
|
--field-separator=::
|
|
|
|
|
|
|
|
Use a special separator character and don't pad with spaces, replacing
|
2010-12-01 09:57:17 +08:00
|
|
|
all occurrences of this separator in symbol names (and other output)
|
2009-07-11 09:47:28 +08:00
|
|
|
with a '.' character, that thus it's the only non valid separator.
|
|
|
|
|
2010-12-01 09:57:17 +08:00
|
|
|
-D::
|
|
|
|
--dump-raw-trace::
|
|
|
|
Dump raw trace in ASCII.
|
|
|
|
|
2011-12-12 23:16:50 +08:00
|
|
|
-g [type,min[,limit],order]::
|
2009-08-31 09:32:03 +08:00
|
|
|
--call-graph::
|
2011-12-12 23:16:50 +08:00
|
|
|
Display call chains using type, min percent threshold, optional print
|
|
|
|
limit and order.
|
2009-08-31 09:32:03 +08:00
|
|
|
type can be either:
|
2010-12-01 09:57:17 +08:00
|
|
|
- flat: single column, linear exposure of call chains.
|
2009-08-31 09:32:03 +08:00
|
|
|
- graph: use a graph tree, displaying absolute overhead rates.
|
|
|
|
- fractal: like graph, but displays relative rates. Each branch of
|
|
|
|
the tree is considered as a new profiled object. +
|
2011-06-07 23:49:46 +08:00
|
|
|
|
|
|
|
order can be either:
|
|
|
|
- callee: callee based call graph.
|
|
|
|
- caller: inverted caller based call graph.
|
|
|
|
|
|
|
|
Default: fractal,0.5,callee.
|
|
|
|
|
|
|
|
-G::
|
|
|
|
--inverted::
|
|
|
|
alias for inverted caller based call graph.
|
2009-08-31 09:32:03 +08:00
|
|
|
|
2010-12-01 09:57:17 +08:00
|
|
|
--pretty=<key>::
|
|
|
|
Pretty printing style. key: normal, raw
|
|
|
|
|
2010-08-21 21:38:16 +08:00
|
|
|
--stdio:: Use the stdio interface.
|
|
|
|
|
|
|
|
--tui:: Use the TUI interface, that is integrated with annotate and allows
|
|
|
|
zooming into DSOs or threads, among other features. Use of --tui
|
|
|
|
requires a tty, if one is not present, as when piping to other
|
|
|
|
commands, the stdio interface is used.
|
|
|
|
|
2010-12-01 09:57:17 +08:00
|
|
|
-k::
|
|
|
|
--vmlinux=<file>::
|
|
|
|
vmlinux pathname
|
|
|
|
|
2010-12-08 10:39:46 +08:00
|
|
|
--kallsyms=<file>::
|
|
|
|
kallsyms pathname
|
|
|
|
|
2010-12-01 09:57:17 +08:00
|
|
|
-m::
|
|
|
|
--modules::
|
|
|
|
Load module symbols. WARNING: This should only be used with -k and
|
|
|
|
a LIVE kernel.
|
|
|
|
|
|
|
|
-f::
|
|
|
|
--force::
|
|
|
|
Don't complain, do it.
|
|
|
|
|
2010-12-10 04:27:07 +08:00
|
|
|
--symfs=<directory>::
|
|
|
|
Look for files with symbols relative to this directory.
|
|
|
|
|
2011-11-14 02:30:08 +08:00
|
|
|
-C::
|
2011-07-04 19:57:50 +08:00
|
|
|
--cpu:: Only report samples for the list of CPUs provided. Multiple CPUs can
|
|
|
|
be provided as a comma-separated list with no space: 0,1. Ranges of
|
|
|
|
CPUs are specified with -: 0-2. Default is to report samples on all
|
|
|
|
CPUs.
|
|
|
|
|
2011-09-16 05:31:41 +08:00
|
|
|
-M::
|
|
|
|
--disassembler-style=:: Set disassembler style for objdump.
|
|
|
|
|
2011-10-06 23:48:31 +08:00
|
|
|
--source::
|
|
|
|
Interleave source code with assembly code. Enabled by default,
|
|
|
|
disable with --no-source.
|
|
|
|
|
|
|
|
--asm-raw::
|
|
|
|
Show raw instruction encoding of assembly instructions.
|
|
|
|
|
2011-10-06 03:10:06 +08:00
|
|
|
--show-total-period:: Show a column with the sum of periods.
|
|
|
|
|
perf tools: Make perf.data more self-descriptive (v8)
The goal of this patch is to include more information about the host
environment into the perf.data so it is more self-descriptive. Overtime,
profiles are captured on various machines and it becomes hard to track
what was recorded, on what machine and when.
This patch provides a way to solve this by extending the perf.data file
with basic information about the host machine. To add those extensions,
we leverage the feature bits capabilities of the perf.data format. The
change is backward compatible with existing perf.data files.
We define the following useful new extensions:
- HEADER_HOSTNAME: the hostname
- HEADER_OSRELEASE: the kernel release number
- HEADER_ARCH: the hw architecture
- HEADER_CPUDESC: generic CPU description
- HEADER_NRCPUS: number of online/avail cpus
- HEADER_CMDLINE: perf command line
- HEADER_VERSION: perf version
- HEADER_TOPOLOGY: cpu topology
- HEADER_EVENT_DESC: full event description (attrs)
- HEADER_CPUID: easy-to-parse low level CPU identication
The small granularity for the entries is to make it easier to extend
without breaking backward compatiblity. Many entries are provided as
ASCII strings.
Perf report/script have been modified to print the basic information as
easy-to-parse ASCII strings. Extended information about CPU and NUMA
topology may be requested with the -I option.
Thanks to David Ahern for reviewing and testing the many versions of
this patch.
$ perf report --stdio
# ========
# captured on : Mon Sep 26 15:22:14 2011
# hostname : quad
# os release : 3.1.0-rc4-tip
# perf version : 3.1.0-rc4
# arch : x86_64
# nrcpus online : 4
# nrcpus avail : 4
# cpudesc : Intel(R) Core(TM)2 Quad CPU Q6600 @ 2.40GHz
# cpuid : GenuineIntel,6,15,11
# total memory : 8105360 kB
# cmdline : /home/eranian/perfmon/official/tip/build/tools/perf/perf record date
# event : name = cycles, type = 0, config = 0x0, config1 = 0x0, config2 = 0x0, excl_usr = 0, excl_kern = 0, id = { 29, 30, 31,
# HEADER_CPU_TOPOLOGY info available, use -I to display
# HEADER_NUMA_TOPOLOGY info available, use -I to display
# ========
#
...
$ perf report --stdio -I
# ========
# captured on : Mon Sep 26 15:22:14 2011
# hostname : quad
# os release : 3.1.0-rc4-tip
# perf version : 3.1.0-rc4
# arch : x86_64
# nrcpus online : 4
# nrcpus avail : 4
# cpudesc : Intel(R) Core(TM)2 Quad CPU Q6600 @ 2.40GHz
# cpuid : GenuineIntel,6,15,11
# total memory : 8105360 kB
# cmdline : /home/eranian/perfmon/official/tip/build/tools/perf/perf record date
# event : name = cycles, type = 0, config = 0x0, config1 = 0x0, config2 = 0x0, excl_usr = 0, excl_kern = 0, id = { 29, 30, 31,
# sibling cores : 0-3
# sibling threads : 0
# sibling threads : 1
# sibling threads : 2
# sibling threads : 3
# node0 meminfo : total = 8320608 kB, free = 7571024 kB
# node0 cpu list : 0-3
# ========
#
...
Reviewed-by: David Ahern <dsahern@gmail.com>
Tested-by: David Ahern <dsahern@gmail.com>
Cc: David Ahern <dsahern@gmail.com>
Cc: Ingo Molnar <mingo@elte.hu>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Robert Richter <robert.richter@amd.com>
Cc: Andi Kleen <ak@linux.intel.com>
Link: http://lkml.kernel.org/r/20110930134040.GA5575@quad
Signed-off-by: Stephane Eranian <eranian@google.com>
[ committer notes: Use --show-info in the tools as was in the docs, rename
perf_header_fprintf_info to perf_file_section__fprintf_info, fixup
conflict with f69b64f7 "perf: Support setting the disassembler style" ]
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
2011-09-30 21:40:40 +08:00
|
|
|
-I::
|
|
|
|
--show-info::
|
|
|
|
Display extended information about the perf.data file. This adds
|
|
|
|
information which may be very large and thus may clutter the display.
|
|
|
|
It currently includes: cpu and numa topology of the host system.
|
|
|
|
|
2009-05-26 15:17:18 +08:00
|
|
|
SEE ALSO
|
|
|
|
--------
|
2011-10-06 23:48:31 +08:00
|
|
|
linkperf:perf-stat[1], linkperf:perf-annotate[1]
|