License cleanup: add SPDX GPL-2.0 license identifier to files with no license
Many source files in the tree are missing licensing information, which
makes it harder for compliance tools to determine the correct license.
By default all files without license information are under the default
license of the kernel, which is GPL version 2.
Update the files which contain no license information with the 'GPL-2.0'
SPDX license identifier. The SPDX identifier is a legally binding
shorthand, which can be used instead of the full boiler plate text.
This patch is based on work done by Thomas Gleixner and Kate Stewart and
Philippe Ombredanne.
How this work was done:
Patches were generated and checked against linux-4.14-rc6 for a subset of
the use cases:
- file had no licensing information it it.
- file was a */uapi/* one with no licensing information in it,
- file was a */uapi/* one with existing licensing information,
Further patches will be generated in subsequent months to fix up cases
where non-standard license headers were used, and references to license
had to be inferred by heuristics based on keywords.
The analysis to determine which SPDX License Identifier to be applied to
a file was done in a spreadsheet of side by side results from of the
output of two independent scanners (ScanCode & Windriver) producing SPDX
tag:value files created by Philippe Ombredanne. Philippe prepared the
base worksheet, and did an initial spot review of a few 1000 files.
The 4.13 kernel was the starting point of the analysis with 60,537 files
assessed. Kate Stewart did a file by file comparison of the scanner
results in the spreadsheet to determine which SPDX license identifier(s)
to be applied to the file. She confirmed any determination that was not
immediately clear with lawyers working with the Linux Foundation.
Criteria used to select files for SPDX license identifier tagging was:
- Files considered eligible had to be source code files.
- Make and config files were included as candidates if they contained >5
lines of source
- File already had some variant of a license header in it (even if <5
lines).
All documentation files were explicitly excluded.
The following heuristics were used to determine which SPDX license
identifiers to apply.
- when both scanners couldn't find any license traces, file was
considered to have no license information in it, and the top level
COPYING file license applied.
For non */uapi/* files that summary was:
SPDX license identifier # files
---------------------------------------------------|-------
GPL-2.0 11139
and resulted in the first patch in this series.
If that file was a */uapi/* path one, it was "GPL-2.0 WITH
Linux-syscall-note" otherwise it was "GPL-2.0". Results of that was:
SPDX license identifier # files
---------------------------------------------------|-------
GPL-2.0 WITH Linux-syscall-note 930
and resulted in the second patch in this series.
- if a file had some form of licensing information in it, and was one
of the */uapi/* ones, it was denoted with the Linux-syscall-note if
any GPL family license was found in the file or had no licensing in
it (per prior point). Results summary:
SPDX license identifier # files
---------------------------------------------------|------
GPL-2.0 WITH Linux-syscall-note 270
GPL-2.0+ WITH Linux-syscall-note 169
((GPL-2.0 WITH Linux-syscall-note) OR BSD-2-Clause) 21
((GPL-2.0 WITH Linux-syscall-note) OR BSD-3-Clause) 17
LGPL-2.1+ WITH Linux-syscall-note 15
GPL-1.0+ WITH Linux-syscall-note 14
((GPL-2.0+ WITH Linux-syscall-note) OR BSD-3-Clause) 5
LGPL-2.0+ WITH Linux-syscall-note 4
LGPL-2.1 WITH Linux-syscall-note 3
((GPL-2.0 WITH Linux-syscall-note) OR MIT) 3
((GPL-2.0 WITH Linux-syscall-note) AND MIT) 1
and that resulted in the third patch in this series.
- when the two scanners agreed on the detected license(s), that became
the concluded license(s).
- when there was disagreement between the two scanners (one detected a
license but the other didn't, or they both detected different
licenses) a manual inspection of the file occurred.
- In most cases a manual inspection of the information in the file
resulted in a clear resolution of the license that should apply (and
which scanner probably needed to revisit its heuristics).
- When it was not immediately clear, the license identifier was
confirmed with lawyers working with the Linux Foundation.
- If there was any question as to the appropriate license identifier,
the file was flagged for further research and to be revisited later
in time.
In total, over 70 hours of logged manual review was done on the
spreadsheet to determine the SPDX license identifiers to apply to the
source files by Kate, Philippe, Thomas and, in some cases, confirmation
by lawyers working with the Linux Foundation.
Kate also obtained a third independent scan of the 4.13 code base from
FOSSology, and compared selected files where the other two scanners
disagreed against that SPDX file, to see if there was new insights. The
Windriver scanner is based on an older version of FOSSology in part, so
they are related.
Thomas did random spot checks in about 500 files from the spreadsheets
for the uapi headers and agreed with SPDX license identifier in the
files he inspected. For the non-uapi files Thomas did random spot checks
in about 15000 files.
In initial set of patches against 4.14-rc6, 3 files were found to have
copy/paste license identifier errors, and have been fixed to reflect the
correct identifier.
Additionally Philippe spent 10 hours this week doing a detailed manual
inspection and review of the 12,461 patched files from the initial patch
version early this week with:
- a full scancode scan run, collecting the matched texts, detected
license ids and scores
- reviewing anything where there was a license detected (about 500+
files) to ensure that the applied SPDX license was correct
- reviewing anything where there was no detection but the patch license
was not GPL-2.0 WITH Linux-syscall-note to ensure that the applied
SPDX license was correct
This produced a worksheet with 20 files needing minor correction. This
worksheet was then exported into 3 different .csv files for the
different types of files to be modified.
These .csv files were then reviewed by Greg. Thomas wrote a script to
parse the csv files and add the proper SPDX tag to the file, in the
format that the file expected. This script was further refined by Greg
based on the output to detect more types of files automatically and to
distinguish between header and source .c files (which need different
comment types.) Finally Greg ran the script using the .csv files to
generate the patches.
Reviewed-by: Kate Stewart <kstewart@linuxfoundation.org>
Reviewed-by: Philippe Ombredanne <pombredanne@nexb.com>
Reviewed-by: Thomas Gleixner <tglx@linutronix.de>
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
2017-11-01 22:07:57 +08:00
|
|
|
/* SPDX-License-Identifier: GPL-2.0 */
|
2008-10-23 13:26:29 +08:00
|
|
|
#ifndef _ASM_X86_PARAVIRT_H
|
|
|
|
#define _ASM_X86_PARAVIRT_H
|
2006-12-07 09:14:07 +08:00
|
|
|
/* Various instructions on x86 need to be replaced for
|
|
|
|
* para-virtualization: those hooks are defined here. */
|
[PATCH] i386: PARAVIRT: Hooks to set up initial pagetable
This patch introduces paravirt_ops hooks to control how the kernel's
initial pagetable is set up.
In the case of a native boot, the very early bootstrap code creates a
simple non-PAE pagetable to map the kernel and physical memory. When
the VM subsystem is initialized, it creates a proper pagetable which
respects the PAE mode, large pages, etc.
When booting under a hypervisor, there are many possibilities for what
paging environment the hypervisor establishes for the guest kernel, so
the constructon of the kernel's pagetable depends on the hypervisor.
In the case of Xen, the hypervisor boots the kernel with a fully
constructed pagetable, which is already using PAE if necessary. Also,
Xen requires particular care when constructing pagetables to make sure
all pagetables are always mapped read-only.
In order to make this easier, kernel's initial pagetable construction
has been changed to only allocate and initialize a pagetable page if
there's no page already present in the pagetable. This allows the Xen
paravirt backend to make a copy of the hypervisor-provided pagetable,
allowing the kernel to establish any more mappings it needs while
keeping the existing ones.
A slightly subtle point which is worth highlighting here is that Xen
requires all kernel mappings to share the same pte_t pages between all
pagetables, so that updating a kernel page's mapping in one pagetable
is reflected in all other pagetables. This makes it possible to
allocate a page and attach it to a pagetable without having to
explicitly enumerate that page's mapping in all pagetables.
And:
+From: "Eric W. Biederman" <ebiederm@xmission.com>
If we don't set the leaf page table entries it is quite possible that
will inherit and incorrect page table entry from the initial boot
page table setup in head.S. So we need to redo the effort here,
so we pick up PSE, PGE and the like.
Hypervisors like Xen require that their page tables be read-only,
which is slightly incompatible with our low identity mappings, however
I discussed this with Jeremy he has modified the Xen early set_pte
function to avoid problems in this area.
Signed-off-by: Eric W. Biederman <ebiederm@xmission.com>
Signed-off-by: Jeremy Fitzhardinge <jeremy@xensource.com>
Signed-off-by: Andi Kleen <ak@suse.de>
Acked-by: William Irwin <bill.irwin@oracle.com>
Cc: Ingo Molnar <mingo@elte.hu>
2007-05-03 01:27:13 +08:00
|
|
|
|
|
|
|
#ifdef CONFIG_PARAVIRT
|
2009-02-12 02:20:05 +08:00
|
|
|
#include <asm/pgtable_types.h>
|
2008-01-30 20:32:06 +08:00
|
|
|
#include <asm/asm.h>
|
2018-01-17 23:58:11 +08:00
|
|
|
#include <asm/nospec-branch.h>
|
2006-12-07 09:14:07 +08:00
|
|
|
|
2009-04-15 05:29:44 +08:00
|
|
|
#include <asm/paravirt_types.h>
|
x86/paravirt: add register-saving thunks to reduce caller register pressure
Impact: Optimization
One of the problems with inserting a pile of C calls where previously
there were none is that the register pressure is greatly increased.
The C calling convention says that the caller must expect a certain
set of registers may be trashed by the callee, and that the callee can
use those registers without restriction. This includes the function
argument registers, and several others.
This patch seeks to alleviate this pressure by introducing wrapper
thunks that will do the register saving/restoring, so that the
callsite doesn't need to worry about it, but the callee function can
be conventional compiler-generated code. In many cases (particularly
performance-sensitive cases) the callee will be in assembler anyway,
and need not use the compiler's calling convention.
Standard calling convention is:
arguments return scratch
x86-32 eax edx ecx eax ?
x86-64 rdi rsi rdx rcx rax r8 r9 r10 r11
The thunk preserves all argument and scratch registers. The return
register is not preserved, and is available as a scratch register for
unwrapped callee code (and of course the return value).
Wrapped function pointers are themselves wrapped in a struct
paravirt_callee_save structure, in order to get some warning from the
compiler when functions with mismatched calling conventions are used.
The most common paravirt ops, both statically and dynamically, are
interrupt enable/disable/save/restore, so handle them first. This is
particularly easy since their calls are handled specially anyway.
XXX Deal with VMI. What's their calling convention?
Signed-off-by: H. Peter Anvin <hpa@zytor.com>
2009-01-29 06:35:05 +08:00
|
|
|
|
2006-12-07 09:14:07 +08:00
|
|
|
#ifndef __ASSEMBLY__
|
2011-11-24 09:12:59 +08:00
|
|
|
#include <linux/bug.h>
|
2007-05-03 01:27:13 +08:00
|
|
|
#include <linux/types.h>
|
2007-05-03 01:27:15 +08:00
|
|
|
#include <linux/cpumask.h>
|
2016-01-22 06:49:13 +08:00
|
|
|
#include <asm/frame.h>
|
2007-05-03 01:27:15 +08:00
|
|
|
|
2018-08-28 15:40:25 +08:00
|
|
|
static inline unsigned long long paravirt_sched_clock(void)
|
|
|
|
{
|
|
|
|
return PVOP_CALL0(unsigned long long, time.sched_clock);
|
|
|
|
}
|
|
|
|
|
|
|
|
struct static_key;
|
|
|
|
extern struct static_key paravirt_steal_enabled;
|
|
|
|
extern struct static_key paravirt_steal_rq_enabled;
|
|
|
|
|
|
|
|
static inline u64 paravirt_steal_clock(int cpu)
|
|
|
|
{
|
|
|
|
return PVOP_CALL1(u64, time.steal_clock, cpu);
|
|
|
|
}
|
|
|
|
|
|
|
|
/* The paravirtualized I/O functions */
|
|
|
|
static inline void slow_down_io(void)
|
|
|
|
{
|
|
|
|
pv_ops.cpu.io_delay();
|
|
|
|
#ifdef REALLY_SLOW_IO
|
|
|
|
pv_ops.cpu.io_delay();
|
|
|
|
pv_ops.cpu.io_delay();
|
|
|
|
pv_ops.cpu.io_delay();
|
|
|
|
#endif
|
|
|
|
}
|
|
|
|
|
|
|
|
static inline void __flush_tlb(void)
|
|
|
|
{
|
|
|
|
PVOP_VCALL0(mmu.flush_tlb_user);
|
|
|
|
}
|
|
|
|
|
|
|
|
static inline void __flush_tlb_global(void)
|
|
|
|
{
|
|
|
|
PVOP_VCALL0(mmu.flush_tlb_kernel);
|
|
|
|
}
|
|
|
|
|
|
|
|
static inline void __flush_tlb_one_user(unsigned long addr)
|
|
|
|
{
|
|
|
|
PVOP_VCALL1(mmu.flush_tlb_one_user, addr);
|
|
|
|
}
|
|
|
|
|
|
|
|
static inline void flush_tlb_others(const struct cpumask *cpumask,
|
|
|
|
const struct flush_tlb_info *info)
|
|
|
|
{
|
|
|
|
PVOP_VCALL2(mmu.flush_tlb_others, cpumask, info);
|
|
|
|
}
|
|
|
|
|
|
|
|
static inline void paravirt_tlb_remove_table(struct mmu_gather *tlb, void *table)
|
|
|
|
{
|
|
|
|
PVOP_VCALL2(mmu.tlb_remove_table, tlb, table);
|
|
|
|
}
|
|
|
|
|
|
|
|
static inline void paravirt_arch_exit_mmap(struct mm_struct *mm)
|
|
|
|
{
|
|
|
|
PVOP_VCALL1(mmu.exit_mmap, mm);
|
|
|
|
}
|
|
|
|
|
2018-08-28 15:40:23 +08:00
|
|
|
#ifdef CONFIG_PARAVIRT_XXL
|
2017-11-02 15:59:10 +08:00
|
|
|
static inline void load_sp0(unsigned long sp0)
|
2006-12-07 09:14:07 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL1(cpu.load_sp0, sp0);
|
2006-12-07 09:14:07 +08:00
|
|
|
}
|
|
|
|
|
|
|
|
/* The paravirtualized CPUID instruction. */
|
|
|
|
static inline void __cpuid(unsigned int *eax, unsigned int *ebx,
|
|
|
|
unsigned int *ecx, unsigned int *edx)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL4(cpu.cpuid, eax, ebx, ecx, edx);
|
2006-12-07 09:14:07 +08:00
|
|
|
}
|
|
|
|
|
|
|
|
/*
|
|
|
|
* These special macros can be used to get or set a debugging register
|
|
|
|
*/
|
2007-05-03 01:27:14 +08:00
|
|
|
static inline unsigned long paravirt_get_debugreg(int reg)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
return PVOP_CALL1(unsigned long, cpu.get_debugreg, reg);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
|
|
|
#define get_debugreg(var, reg) var = paravirt_get_debugreg(reg)
|
|
|
|
static inline void set_debugreg(unsigned long val, int reg)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL2(cpu.set_debugreg, reg, val);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
2006-12-07 09:14:07 +08:00
|
|
|
|
2007-05-03 01:27:14 +08:00
|
|
|
static inline unsigned long read_cr0(void)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
return PVOP_CALL0(unsigned long, cpu.read_cr0);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
2006-12-07 09:14:07 +08:00
|
|
|
|
2007-05-03 01:27:14 +08:00
|
|
|
static inline void write_cr0(unsigned long x)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL1(cpu.write_cr0, x);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
|
|
|
|
|
|
|
static inline unsigned long read_cr2(void)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
return PVOP_CALL0(unsigned long, mmu.read_cr2);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
|
|
|
|
|
|
|
static inline void write_cr2(unsigned long x)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL1(mmu.write_cr2, x);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
|
|
|
|
2017-06-13 01:26:14 +08:00
|
|
|
static inline unsigned long __read_cr3(void)
|
2007-05-03 01:27:14 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
return PVOP_CALL0(unsigned long, mmu.read_cr3);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
2006-12-07 09:14:07 +08:00
|
|
|
|
2007-05-03 01:27:14 +08:00
|
|
|
static inline void write_cr3(unsigned long x)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL1(mmu.write_cr3, x);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
2006-12-07 09:14:07 +08:00
|
|
|
|
2014-10-25 06:58:08 +08:00
|
|
|
static inline void __write_cr4(unsigned long x)
|
2007-05-03 01:27:14 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL1(cpu.write_cr4, x);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
2007-05-03 01:27:13 +08:00
|
|
|
|
2008-01-30 20:33:19 +08:00
|
|
|
#ifdef CONFIG_X86_64
|
2008-01-30 20:33:19 +08:00
|
|
|
static inline unsigned long read_cr8(void)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
return PVOP_CALL0(unsigned long, cpu.read_cr8);
|
2008-01-30 20:33:19 +08:00
|
|
|
}
|
|
|
|
|
|
|
|
static inline void write_cr8(unsigned long x)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL1(cpu.write_cr8, x);
|
2008-01-30 20:33:19 +08:00
|
|
|
}
|
2008-01-30 20:33:19 +08:00
|
|
|
#endif
|
2008-01-30 20:33:19 +08:00
|
|
|
|
2010-10-07 21:08:55 +08:00
|
|
|
static inline void arch_safe_halt(void)
|
2006-12-07 09:14:07 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL0(irq.safe_halt);
|
2006-12-07 09:14:07 +08:00
|
|
|
}
|
|
|
|
|
|
|
|
static inline void halt(void)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL0(irq.halt);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
|
|
|
|
|
|
|
static inline void wbinvd(void)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL0(cpu.wbinvd);
|
2006-12-07 09:14:07 +08:00
|
|
|
}
|
|
|
|
|
2007-10-17 02:51:29 +08:00
|
|
|
#define get_kernel_rpl() (pv_info.kernel_rpl)
|
2006-12-07 09:14:07 +08:00
|
|
|
|
2016-04-02 22:01:38 +08:00
|
|
|
static inline u64 paravirt_read_msr(unsigned msr)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
return PVOP_CALL1(u64, cpu.read_msr, msr);
|
2016-04-02 22:01:38 +08:00
|
|
|
}
|
|
|
|
|
|
|
|
static inline void paravirt_write_msr(unsigned msr,
|
|
|
|
unsigned low, unsigned high)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL3(cpu.write_msr, msr, low, high);
|
2016-04-02 22:01:38 +08:00
|
|
|
}
|
|
|
|
|
2016-04-02 22:01:36 +08:00
|
|
|
static inline u64 paravirt_read_msr_safe(unsigned msr, int *err)
|
2007-05-03 01:27:14 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
return PVOP_CALL2(u64, cpu.read_msr_safe, msr, err);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
2009-08-31 15:50:09 +08:00
|
|
|
|
2016-04-02 22:01:36 +08:00
|
|
|
static inline int paravirt_write_msr_safe(unsigned msr,
|
|
|
|
unsigned low, unsigned high)
|
2007-05-03 01:27:14 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
return PVOP_CALL3(int, cpu.write_msr_safe, msr, low, high);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
|
|
|
|
2008-03-23 16:03:00 +08:00
|
|
|
#define rdmsr(msr, val1, val2) \
|
|
|
|
do { \
|
2016-04-02 22:01:39 +08:00
|
|
|
u64 _l = paravirt_read_msr(msr); \
|
2007-05-03 01:27:14 +08:00
|
|
|
val1 = (u32)_l; \
|
|
|
|
val2 = _l >> 32; \
|
2008-03-23 16:03:00 +08:00
|
|
|
} while (0)
|
2006-12-07 09:14:07 +08:00
|
|
|
|
2008-03-23 16:03:00 +08:00
|
|
|
#define wrmsr(msr, val1, val2) \
|
|
|
|
do { \
|
2016-04-02 22:01:39 +08:00
|
|
|
paravirt_write_msr(msr, val1, val2); \
|
2008-03-23 16:03:00 +08:00
|
|
|
} while (0)
|
2006-12-07 09:14:07 +08:00
|
|
|
|
2008-03-23 16:03:00 +08:00
|
|
|
#define rdmsrl(msr, val) \
|
|
|
|
do { \
|
2016-04-02 22:01:39 +08:00
|
|
|
val = paravirt_read_msr(msr); \
|
2008-03-23 16:03:00 +08:00
|
|
|
} while (0)
|
2006-12-07 09:14:07 +08:00
|
|
|
|
2015-07-24 03:14:40 +08:00
|
|
|
static inline void wrmsrl(unsigned msr, u64 val)
|
|
|
|
{
|
|
|
|
wrmsr(msr, (u32)val, (u32)(val>>32));
|
|
|
|
}
|
|
|
|
|
2016-04-02 22:01:36 +08:00
|
|
|
#define wrmsr_safe(msr, a, b) paravirt_write_msr_safe(msr, a, b)
|
2006-12-07 09:14:07 +08:00
|
|
|
|
|
|
|
/* rdmsr with exception handling */
|
2016-04-02 22:01:36 +08:00
|
|
|
#define rdmsr_safe(msr, a, b) \
|
|
|
|
({ \
|
|
|
|
int _err; \
|
|
|
|
u64 _l = paravirt_read_msr_safe(msr, &_err); \
|
|
|
|
(*a) = (u32)_l; \
|
|
|
|
(*b) = _l >> 32; \
|
|
|
|
_err; \
|
2008-03-23 16:03:00 +08:00
|
|
|
})
|
2006-12-07 09:14:07 +08:00
|
|
|
|
2008-03-22 17:59:28 +08:00
|
|
|
static inline int rdmsrl_safe(unsigned msr, unsigned long long *p)
|
|
|
|
{
|
|
|
|
int err;
|
|
|
|
|
2016-04-02 22:01:36 +08:00
|
|
|
*p = paravirt_read_msr_safe(msr, &err);
|
2008-03-22 17:59:28 +08:00
|
|
|
return err;
|
|
|
|
}
|
2009-08-31 15:50:10 +08:00
|
|
|
|
2007-05-03 01:27:14 +08:00
|
|
|
static inline unsigned long long paravirt_read_pmc(int counter)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
return PVOP_CALL1(u64, cpu.read_pmc, counter);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
2006-12-07 09:14:07 +08:00
|
|
|
|
2008-03-23 16:03:00 +08:00
|
|
|
#define rdpmc(counter, low, high) \
|
|
|
|
do { \
|
2007-05-03 01:27:14 +08:00
|
|
|
u64 _l = paravirt_read_pmc(counter); \
|
|
|
|
low = (u32)_l; \
|
|
|
|
high = _l >> 32; \
|
2008-03-23 16:03:00 +08:00
|
|
|
} while (0)
|
2007-05-03 01:27:13 +08:00
|
|
|
|
2012-06-06 08:56:50 +08:00
|
|
|
#define rdpmcl(counter, val) ((val) = paravirt_read_pmc(counter))
|
|
|
|
|
2008-07-24 05:21:18 +08:00
|
|
|
static inline void paravirt_alloc_ldt(struct desc_struct *ldt, unsigned entries)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL2(cpu.alloc_ldt, ldt, entries);
|
2008-07-24 05:21:18 +08:00
|
|
|
}
|
|
|
|
|
|
|
|
static inline void paravirt_free_ldt(struct desc_struct *ldt, unsigned entries)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL2(cpu.free_ldt, ldt, entries);
|
2008-07-24 05:21:18 +08:00
|
|
|
}
|
|
|
|
|
2007-05-03 01:27:14 +08:00
|
|
|
static inline void load_TR_desc(void)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL0(cpu.load_tr_desc);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
2008-01-30 20:31:12 +08:00
|
|
|
static inline void load_gdt(const struct desc_ptr *dtr)
|
2007-05-03 01:27:14 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL1(cpu.load_gdt, dtr);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
2008-01-30 20:31:12 +08:00
|
|
|
static inline void load_idt(const struct desc_ptr *dtr)
|
2007-05-03 01:27:14 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL1(cpu.load_idt, dtr);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
|
|
|
static inline void set_ldt(const void *addr, unsigned entries)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL2(cpu.set_ldt, addr, entries);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
|
|
|
static inline unsigned long paravirt_store_tr(void)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
return PVOP_CALL0(unsigned long, cpu.store_tr);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
2018-08-28 15:40:23 +08:00
|
|
|
|
2007-05-03 01:27:14 +08:00
|
|
|
#define store_tr(tr) ((tr) = paravirt_store_tr())
|
|
|
|
static inline void load_TLS(struct thread_struct *t, unsigned cpu)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL2(cpu.load_tls, t, cpu);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
2008-01-30 20:31:13 +08:00
|
|
|
|
2008-06-25 12:19:32 +08:00
|
|
|
#ifdef CONFIG_X86_64
|
|
|
|
static inline void load_gs_index(unsigned int gs)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL1(cpu.load_gs_index, gs);
|
2008-06-25 12:19:32 +08:00
|
|
|
}
|
|
|
|
#endif
|
|
|
|
|
2008-01-30 20:31:13 +08:00
|
|
|
static inline void write_ldt_entry(struct desc_struct *dt, int entry,
|
|
|
|
const void *desc)
|
2007-05-03 01:27:14 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL3(cpu.write_ldt_entry, dt, entry, desc);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
2008-01-30 20:31:13 +08:00
|
|
|
|
|
|
|
static inline void write_gdt_entry(struct desc_struct *dt, int entry,
|
|
|
|
void *desc, int type)
|
2007-05-03 01:27:14 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL4(cpu.write_gdt_entry, dt, entry, desc, type);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
2008-01-30 20:31:13 +08:00
|
|
|
|
2008-01-30 20:31:12 +08:00
|
|
|
static inline void write_idt_entry(gate_desc *dt, int entry, const gate_desc *g)
|
2007-05-03 01:27:14 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL3(cpu.write_idt_entry, dt, entry, g);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
|
|
|
static inline void set_iopl_mask(unsigned mask)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL1(cpu.set_iopl_mask, mask);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
2006-12-07 09:14:07 +08:00
|
|
|
|
2007-05-03 01:27:14 +08:00
|
|
|
static inline void paravirt_activate_mm(struct mm_struct *prev,
|
|
|
|
struct mm_struct *next)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL2(mmu.activate_mm, prev, next);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
|
|
|
|
2014-11-19 02:23:49 +08:00
|
|
|
static inline void paravirt_arch_dup_mmap(struct mm_struct *oldmm,
|
|
|
|
struct mm_struct *mm)
|
2007-05-03 01:27:14 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL2(mmu.dup_mmap, oldmm, mm);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
|
|
|
|
2008-06-25 12:19:12 +08:00
|
|
|
static inline int paravirt_pgd_alloc(struct mm_struct *mm)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
return PVOP_CALL1(int, mmu.pgd_alloc, mm);
|
2008-06-25 12:19:12 +08:00
|
|
|
}
|
|
|
|
|
|
|
|
static inline void paravirt_pgd_free(struct mm_struct *mm, pgd_t *pgd)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL2(mmu.pgd_free, mm, pgd);
|
2008-06-25 12:19:12 +08:00
|
|
|
}
|
|
|
|
|
2008-07-31 05:32:27 +08:00
|
|
|
static inline void paravirt_alloc_pte(struct mm_struct *mm, unsigned long pfn)
|
2007-05-03 01:27:14 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL2(mmu.alloc_pte, mm, pfn);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
2008-07-31 05:32:27 +08:00
|
|
|
static inline void paravirt_release_pte(unsigned long pfn)
|
2007-05-03 01:27:14 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL1(mmu.release_pte, pfn);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
2007-02-13 20:26:21 +08:00
|
|
|
|
2008-07-31 05:32:27 +08:00
|
|
|
static inline void paravirt_alloc_pmd(struct mm_struct *mm, unsigned long pfn)
|
2007-05-03 01:27:14 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL2(mmu.alloc_pmd, mm, pfn);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
2007-02-13 20:26:21 +08:00
|
|
|
|
2008-07-31 05:32:27 +08:00
|
|
|
static inline void paravirt_release_pmd(unsigned long pfn)
|
2006-12-07 09:14:08 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL1(mmu.release_pmd, pfn);
|
2006-12-07 09:14:08 +08:00
|
|
|
}
|
|
|
|
|
2008-07-31 05:32:27 +08:00
|
|
|
static inline void paravirt_alloc_pud(struct mm_struct *mm, unsigned long pfn)
|
2008-03-18 07:37:02 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL2(mmu.alloc_pud, mm, pfn);
|
2008-03-18 07:37:02 +08:00
|
|
|
}
|
2008-07-31 05:32:27 +08:00
|
|
|
static inline void paravirt_release_pud(unsigned long pfn)
|
2008-03-18 07:37:02 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL1(mmu.release_pud, pfn);
|
2008-03-18 07:37:02 +08:00
|
|
|
}
|
|
|
|
|
2017-03-30 16:07:28 +08:00
|
|
|
static inline void paravirt_alloc_p4d(struct mm_struct *mm, unsigned long pfn)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL2(mmu.alloc_p4d, mm, pfn);
|
2017-03-30 16:07:28 +08:00
|
|
|
}
|
|
|
|
|
|
|
|
static inline void paravirt_release_p4d(unsigned long pfn)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL1(mmu.release_p4d, pfn);
|
2017-03-30 16:07:28 +08:00
|
|
|
}
|
|
|
|
|
2008-01-30 20:33:15 +08:00
|
|
|
static inline pte_t __pte(pteval_t val)
|
2006-12-07 09:14:08 +08:00
|
|
|
{
|
2008-01-30 20:33:15 +08:00
|
|
|
pteval_t ret;
|
|
|
|
|
|
|
|
if (sizeof(pteval_t) > sizeof(long))
|
2018-08-28 15:40:19 +08:00
|
|
|
ret = PVOP_CALLEE2(pteval_t, mmu.make_pte, val, (u64)val >> 32);
|
2008-01-30 20:33:15 +08:00
|
|
|
else
|
2018-08-28 15:40:19 +08:00
|
|
|
ret = PVOP_CALLEE1(pteval_t, mmu.make_pte, val);
|
2008-01-30 20:33:15 +08:00
|
|
|
|
2008-01-30 20:32:57 +08:00
|
|
|
return (pte_t) { .pte = ret };
|
2006-12-07 09:14:08 +08:00
|
|
|
}
|
|
|
|
|
2008-01-30 20:33:15 +08:00
|
|
|
static inline pteval_t pte_val(pte_t pte)
|
|
|
|
{
|
|
|
|
pteval_t ret;
|
|
|
|
|
|
|
|
if (sizeof(pteval_t) > sizeof(long))
|
2018-08-28 15:40:19 +08:00
|
|
|
ret = PVOP_CALLEE2(pteval_t, mmu.pte_val,
|
2009-01-29 06:35:07 +08:00
|
|
|
pte.pte, (u64)pte.pte >> 32);
|
2008-01-30 20:33:15 +08:00
|
|
|
else
|
2018-08-28 15:40:19 +08:00
|
|
|
ret = PVOP_CALLEE1(pteval_t, mmu.pte_val, pte.pte);
|
2008-01-30 20:33:15 +08:00
|
|
|
|
|
|
|
return ret;
|
|
|
|
}
|
|
|
|
|
2008-01-30 20:33:15 +08:00
|
|
|
static inline pgd_t __pgd(pgdval_t val)
|
2006-12-07 09:14:08 +08:00
|
|
|
{
|
2008-01-30 20:33:15 +08:00
|
|
|
pgdval_t ret;
|
|
|
|
|
|
|
|
if (sizeof(pgdval_t) > sizeof(long))
|
2018-08-28 15:40:19 +08:00
|
|
|
ret = PVOP_CALLEE2(pgdval_t, mmu.make_pgd, val, (u64)val >> 32);
|
2008-01-30 20:33:15 +08:00
|
|
|
else
|
2018-08-28 15:40:19 +08:00
|
|
|
ret = PVOP_CALLEE1(pgdval_t, mmu.make_pgd, val);
|
2008-01-30 20:33:15 +08:00
|
|
|
|
|
|
|
return (pgd_t) { ret };
|
|
|
|
}
|
|
|
|
|
|
|
|
static inline pgdval_t pgd_val(pgd_t pgd)
|
|
|
|
{
|
|
|
|
pgdval_t ret;
|
|
|
|
|
|
|
|
if (sizeof(pgdval_t) > sizeof(long))
|
2018-08-28 15:40:19 +08:00
|
|
|
ret = PVOP_CALLEE2(pgdval_t, mmu.pgd_val,
|
2009-01-29 06:35:07 +08:00
|
|
|
pgd.pgd, (u64)pgd.pgd >> 32);
|
2008-01-30 20:33:15 +08:00
|
|
|
else
|
2018-08-28 15:40:19 +08:00
|
|
|
ret = PVOP_CALLEE1(pgdval_t, mmu.pgd_val, pgd.pgd);
|
2008-01-30 20:33:15 +08:00
|
|
|
|
|
|
|
return ret;
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
|
|
|
|
2008-06-16 19:30:01 +08:00
|
|
|
#define __HAVE_ARCH_PTEP_MODIFY_PROT_TRANSACTION
|
|
|
|
static inline pte_t ptep_modify_prot_start(struct mm_struct *mm, unsigned long addr,
|
|
|
|
pte_t *ptep)
|
|
|
|
{
|
|
|
|
pteval_t ret;
|
|
|
|
|
2018-08-28 15:40:19 +08:00
|
|
|
ret = PVOP_CALL3(pteval_t, mmu.ptep_modify_prot_start, mm, addr, ptep);
|
2008-06-16 19:30:01 +08:00
|
|
|
|
|
|
|
return (pte_t) { .pte = ret };
|
|
|
|
}
|
|
|
|
|
|
|
|
static inline void ptep_modify_prot_commit(struct mm_struct *mm, unsigned long addr,
|
|
|
|
pte_t *ptep, pte_t pte)
|
|
|
|
{
|
|
|
|
if (sizeof(pteval_t) > sizeof(long))
|
|
|
|
/* 5 arg words */
|
2018-08-28 15:40:19 +08:00
|
|
|
pv_ops.mmu.ptep_modify_prot_commit(mm, addr, ptep, pte);
|
2008-06-16 19:30:01 +08:00
|
|
|
else
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL4(mmu.ptep_modify_prot_commit,
|
2008-06-16 19:30:01 +08:00
|
|
|
mm, addr, ptep, pte.pte);
|
|
|
|
}
|
|
|
|
|
2008-01-30 20:33:15 +08:00
|
|
|
static inline void set_pte(pte_t *ptep, pte_t pte)
|
|
|
|
{
|
|
|
|
if (sizeof(pteval_t) > sizeof(long))
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL3(mmu.set_pte, ptep, pte.pte, (u64)pte.pte >> 32);
|
2008-01-30 20:33:15 +08:00
|
|
|
else
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL2(mmu.set_pte, ptep, pte.pte);
|
2008-01-30 20:33:15 +08:00
|
|
|
}
|
|
|
|
|
|
|
|
static inline void set_pte_at(struct mm_struct *mm, unsigned long addr,
|
|
|
|
pte_t *ptep, pte_t pte)
|
|
|
|
{
|
|
|
|
if (sizeof(pteval_t) > sizeof(long))
|
|
|
|
/* 5 arg words */
|
2018-08-28 15:40:19 +08:00
|
|
|
pv_ops.mmu.set_pte_at(mm, addr, ptep, pte);
|
2008-01-30 20:33:15 +08:00
|
|
|
else
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL4(mmu.set_pte_at, mm, addr, ptep, pte.pte);
|
2008-01-30 20:33:15 +08:00
|
|
|
}
|
|
|
|
|
2008-01-30 20:33:15 +08:00
|
|
|
static inline void set_pmd(pmd_t *pmdp, pmd_t pmd)
|
|
|
|
{
|
|
|
|
pmdval_t val = native_pmd_val(pmd);
|
|
|
|
|
|
|
|
if (sizeof(pmdval_t) > sizeof(long))
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL3(mmu.set_pmd, pmdp, val, (u64)val >> 32);
|
2008-01-30 20:33:15 +08:00
|
|
|
else
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL2(mmu.set_pmd, pmdp, val);
|
2008-01-30 20:33:15 +08:00
|
|
|
}
|
|
|
|
|
2015-04-15 06:46:14 +08:00
|
|
|
#if CONFIG_PGTABLE_LEVELS >= 3
|
2008-01-30 20:33:19 +08:00
|
|
|
static inline pmd_t __pmd(pmdval_t val)
|
|
|
|
{
|
|
|
|
pmdval_t ret;
|
|
|
|
|
|
|
|
if (sizeof(pmdval_t) > sizeof(long))
|
2018-08-28 15:40:19 +08:00
|
|
|
ret = PVOP_CALLEE2(pmdval_t, mmu.make_pmd, val, (u64)val >> 32);
|
2008-01-30 20:33:19 +08:00
|
|
|
else
|
2018-08-28 15:40:19 +08:00
|
|
|
ret = PVOP_CALLEE1(pmdval_t, mmu.make_pmd, val);
|
2008-01-30 20:33:19 +08:00
|
|
|
|
|
|
|
return (pmd_t) { ret };
|
|
|
|
}
|
|
|
|
|
|
|
|
static inline pmdval_t pmd_val(pmd_t pmd)
|
|
|
|
{
|
|
|
|
pmdval_t ret;
|
|
|
|
|
|
|
|
if (sizeof(pmdval_t) > sizeof(long))
|
2018-08-28 15:40:19 +08:00
|
|
|
ret = PVOP_CALLEE2(pmdval_t, mmu.pmd_val,
|
2009-01-29 06:35:07 +08:00
|
|
|
pmd.pmd, (u64)pmd.pmd >> 32);
|
2008-01-30 20:33:19 +08:00
|
|
|
else
|
2018-08-28 15:40:19 +08:00
|
|
|
ret = PVOP_CALLEE1(pmdval_t, mmu.pmd_val, pmd.pmd);
|
2008-01-30 20:33:19 +08:00
|
|
|
|
|
|
|
return ret;
|
|
|
|
}
|
|
|
|
|
|
|
|
static inline void set_pud(pud_t *pudp, pud_t pud)
|
|
|
|
{
|
|
|
|
pudval_t val = native_pud_val(pud);
|
|
|
|
|
|
|
|
if (sizeof(pudval_t) > sizeof(long))
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL3(mmu.set_pud, pudp, val, (u64)val >> 32);
|
2008-01-30 20:33:19 +08:00
|
|
|
else
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL2(mmu.set_pud, pudp, val);
|
2008-01-30 20:33:19 +08:00
|
|
|
}
|
2017-03-18 02:55:15 +08:00
|
|
|
#if CONFIG_PGTABLE_LEVELS >= 4
|
2008-01-30 20:33:20 +08:00
|
|
|
static inline pud_t __pud(pudval_t val)
|
|
|
|
{
|
|
|
|
pudval_t ret;
|
|
|
|
|
2018-08-28 15:40:26 +08:00
|
|
|
ret = PVOP_CALLEE1(pudval_t, mmu.make_pud, val);
|
2008-01-30 20:33:20 +08:00
|
|
|
|
|
|
|
return (pud_t) { ret };
|
|
|
|
}
|
|
|
|
|
|
|
|
static inline pudval_t pud_val(pud_t pud)
|
|
|
|
{
|
2018-08-28 15:40:26 +08:00
|
|
|
return PVOP_CALLEE1(pudval_t, mmu.pud_val, pud.pud);
|
2008-01-30 20:33:20 +08:00
|
|
|
}
|
|
|
|
|
2017-03-18 02:55:15 +08:00
|
|
|
static inline void pud_clear(pud_t *pudp)
|
|
|
|
{
|
|
|
|
set_pud(pudp, __pud(0));
|
|
|
|
}
|
|
|
|
|
|
|
|
static inline void set_p4d(p4d_t *p4dp, p4d_t p4d)
|
|
|
|
{
|
|
|
|
p4dval_t val = native_p4d_val(p4d);
|
|
|
|
|
2018-08-28 15:40:26 +08:00
|
|
|
PVOP_VCALL2(mmu.set_p4d, p4dp, val);
|
2017-03-18 02:55:15 +08:00
|
|
|
}
|
|
|
|
|
2017-03-30 16:07:28 +08:00
|
|
|
#if CONFIG_PGTABLE_LEVELS >= 5
|
|
|
|
|
|
|
|
static inline p4d_t __p4d(p4dval_t val)
|
2017-03-18 02:55:15 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
p4dval_t ret = PVOP_CALLEE1(p4dval_t, mmu.make_p4d, val);
|
2017-03-18 02:55:15 +08:00
|
|
|
|
2017-03-30 16:07:28 +08:00
|
|
|
return (p4d_t) { ret };
|
|
|
|
}
|
2017-03-18 02:55:15 +08:00
|
|
|
|
2017-03-30 16:07:28 +08:00
|
|
|
static inline p4dval_t p4d_val(p4d_t p4d)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
return PVOP_CALLEE1(p4dval_t, mmu.p4d_val, p4d.p4d);
|
2017-03-30 16:07:28 +08:00
|
|
|
}
|
2017-03-18 02:55:15 +08:00
|
|
|
|
2018-02-16 19:49:47 +08:00
|
|
|
static inline void __set_pgd(pgd_t *pgdp, pgd_t pgd)
|
2008-01-30 20:33:20 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL2(mmu.set_pgd, pgdp, native_pgd_val(pgd));
|
2008-01-30 20:33:20 +08:00
|
|
|
}
|
|
|
|
|
2018-02-16 19:49:47 +08:00
|
|
|
#define set_pgd(pgdp, pgdval) do { \
|
2018-05-18 18:35:24 +08:00
|
|
|
if (pgtable_l5_enabled()) \
|
2018-02-16 19:49:47 +08:00
|
|
|
__set_pgd(pgdp, pgdval); \
|
|
|
|
else \
|
|
|
|
set_p4d((p4d_t *)(pgdp), (p4d_t) { (pgdval).pgd }); \
|
|
|
|
} while (0)
|
|
|
|
|
|
|
|
#define pgd_clear(pgdp) do { \
|
2018-05-18 18:35:24 +08:00
|
|
|
if (pgtable_l5_enabled()) \
|
2018-02-16 19:49:47 +08:00
|
|
|
set_pgd(pgdp, __pgd(0)); \
|
|
|
|
} while (0)
|
2008-01-30 20:33:20 +08:00
|
|
|
|
2017-03-18 02:55:15 +08:00
|
|
|
#endif /* CONFIG_PGTABLE_LEVELS == 5 */
|
2008-01-30 20:33:20 +08:00
|
|
|
|
2017-03-30 16:07:28 +08:00
|
|
|
static inline void p4d_clear(p4d_t *p4dp)
|
|
|
|
{
|
|
|
|
set_p4d(p4dp, __p4d(0));
|
|
|
|
}
|
|
|
|
|
2015-04-15 06:46:14 +08:00
|
|
|
#endif /* CONFIG_PGTABLE_LEVELS == 4 */
|
2008-01-30 20:33:20 +08:00
|
|
|
|
2015-04-15 06:46:14 +08:00
|
|
|
#endif /* CONFIG_PGTABLE_LEVELS >= 3 */
|
2008-01-30 20:33:19 +08:00
|
|
|
|
2008-01-30 20:33:15 +08:00
|
|
|
#ifdef CONFIG_X86_PAE
|
|
|
|
/* Special-case pte-setting operations for PAE, which can't update a
|
|
|
|
64-bit pte atomically */
|
|
|
|
static inline void set_pte_atomic(pte_t *ptep, pte_t pte)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL3(mmu.set_pte_atomic, ptep, pte.pte, pte.pte >> 32);
|
2008-01-30 20:33:15 +08:00
|
|
|
}
|
|
|
|
|
|
|
|
static inline void pte_clear(struct mm_struct *mm, unsigned long addr,
|
|
|
|
pte_t *ptep)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL3(mmu.pte_clear, mm, addr, ptep);
|
2008-01-30 20:33:15 +08:00
|
|
|
}
|
2008-01-30 20:33:15 +08:00
|
|
|
|
|
|
|
static inline void pmd_clear(pmd_t *pmdp)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL1(mmu.pmd_clear, pmdp);
|
2008-01-30 20:33:15 +08:00
|
|
|
}
|
2008-01-30 20:33:15 +08:00
|
|
|
#else /* !CONFIG_X86_PAE */
|
|
|
|
static inline void set_pte_atomic(pte_t *ptep, pte_t pte)
|
|
|
|
{
|
|
|
|
set_pte(ptep, pte);
|
|
|
|
}
|
|
|
|
|
|
|
|
static inline void pte_clear(struct mm_struct *mm, unsigned long addr,
|
|
|
|
pte_t *ptep)
|
|
|
|
{
|
|
|
|
set_pte_at(mm, addr, ptep, __pte(0));
|
|
|
|
}
|
2008-01-30 20:33:15 +08:00
|
|
|
|
|
|
|
static inline void pmd_clear(pmd_t *pmdp)
|
|
|
|
{
|
|
|
|
set_pmd(pmdp, __pmd(0));
|
|
|
|
}
|
2008-01-30 20:33:15 +08:00
|
|
|
#endif /* CONFIG_X86_PAE */
|
|
|
|
|
2009-02-18 15:24:03 +08:00
|
|
|
#define __HAVE_ARCH_START_CONTEXT_SWITCH
|
2009-02-19 03:18:57 +08:00
|
|
|
static inline void arch_start_context_switch(struct task_struct *prev)
|
2007-05-03 01:27:14 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL1(cpu.start_context_switch, prev);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
|
|
|
|
2009-02-19 03:18:57 +08:00
|
|
|
static inline void arch_end_context_switch(struct task_struct *next)
|
2007-05-03 01:27:14 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL1(cpu.end_context_switch, next);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
|
|
|
|
2007-02-13 20:26:21 +08:00
|
|
|
#define __HAVE_ARCH_ENTER_LAZY_MMU_MODE
|
2007-05-03 01:27:14 +08:00
|
|
|
static inline void arch_enter_lazy_mmu_mode(void)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL0(mmu.lazy_mode.enter);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
|
|
|
|
|
|
|
static inline void arch_leave_lazy_mmu_mode(void)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL0(mmu.lazy_mode.leave);
|
2007-05-03 01:27:14 +08:00
|
|
|
}
|
|
|
|
|
2013-03-23 21:36:36 +08:00
|
|
|
static inline void arch_flush_lazy_mmu_mode(void)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL0(mmu.lazy_mode.flush);
|
2013-03-23 21:36:36 +08:00
|
|
|
}
|
2007-02-13 20:26:21 +08:00
|
|
|
|
2008-06-18 02:42:01 +08:00
|
|
|
static inline void __set_fixmap(unsigned /* enum fixed_addresses */ idx,
|
2009-04-10 01:55:33 +08:00
|
|
|
phys_addr_t phys, pgprot_t flags)
|
2008-06-18 02:42:01 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
pv_ops.mmu.set_fixmap(idx, phys, flags);
|
2008-06-18 02:42:01 +08:00
|
|
|
}
|
2018-08-28 15:40:25 +08:00
|
|
|
#endif
|
2008-06-18 02:42:01 +08:00
|
|
|
|
x86: Fix performance regression caused by paravirt_ops on native kernels
Xiaohui Xin and some other folks at Intel have been looking into what's
behind the performance hit of paravirt_ops when running native.
It appears that the hit is entirely due to the paravirtualized
spinlocks introduced by:
| commit 8efcbab674de2bee45a2e4cdf97de16b8e609ac8
| Date: Mon Jul 7 12:07:51 2008 -0700
|
| paravirt: introduce a "lock-byte" spinlock implementation
The extra call/return in the spinlock path is somehow
causing an increase in the cycles/instruction of somewhere around 2-7%
(seems to vary quite a lot from test to test). The working theory is
that the CPU's pipeline is getting upset about the
call->call->locked-op->return->return, and seems to be failing to
speculate (though I haven't seen anything definitive about the precise
reasons). This doesn't entirely make sense, because the performance
hit is also visible on unlock and other operations which don't involve
locked instructions. But spinlock operations clearly swamp all the
other pvops operations, even though I can't imagine that they're
nearly as common (there's only a .05% increase in instructions
executed).
If I disable just the pv-spinlock calls, my tests show that pvops is
identical to non-pvops performance on native (my measurements show that
it is actually about .1% faster, but Xiaohui shows a .05% slowdown).
Summary of results, averaging 10 runs of the "mmperf" test, using a
no-pvops build as baseline:
nopv Pv-nospin Pv-spin
CPU cycles 100.00% 99.89% 102.18%
instructions 100.00% 100.10% 100.15%
CPI 100.00% 99.79% 102.03%
cache ref 100.00% 100.84% 100.28%
cache miss 100.00% 90.47% 88.56%
cache miss rate 100.00% 89.72% 88.31%
branches 100.00% 99.93% 100.04%
branch miss 100.00% 103.66% 107.72%
branch miss rt 100.00% 103.73% 107.67%
wallclock 100.00% 99.90% 102.20%
The clear effect here is that the 2% increase in CPI is
directly reflected in the final wallclock time.
(The other interesting effect is that the more ops are
out of line calls via pvops, the lower the cache access
and miss rates. Not too surprising, but it suggests that
the non-pvops kernel is over-inlined. On the flipside,
the branch misses go up correspondingly...)
So, what's the fix?
Paravirt patching turns all the pvops calls into direct calls, so
_spin_lock etc do end up having direct calls. For example, the compiler
generated code for paravirtualized _spin_lock is:
<_spin_lock+0>: mov %gs:0xb4c8,%rax
<_spin_lock+9>: incl 0xffffffffffffe044(%rax)
<_spin_lock+15>: callq *0xffffffff805a5b30
<_spin_lock+22>: retq
The indirect call will get patched to:
<_spin_lock+0>: mov %gs:0xb4c8,%rax
<_spin_lock+9>: incl 0xffffffffffffe044(%rax)
<_spin_lock+15>: callq <__ticket_spin_lock>
<_spin_lock+20>: nop; nop /* or whatever 2-byte nop */
<_spin_lock+22>: retq
One possibility is to inline _spin_lock, etc, when building an
optimised kernel (ie, when there's no spinlock/preempt
instrumentation/debugging enabled). That will remove the outer
call/return pair, returning the instruction stream to a single
call/return, which will presumably execute the same as the non-pvops
case. The downsides arel 1) it will replicate the
preempt_disable/enable code at eack lock/unlock callsite; this code is
fairly small, but not nothing; and 2) the spinlock definitions are
already a very heavily tangled mass of #ifdefs and other preprocessor
magic, and making any changes will be non-trivial.
The other obvious answer is to disable pv-spinlocks. Making them a
separate config option is fairly easy, and it would be trivial to
enable them only when Xen is enabled (as the only non-default user).
But it doesn't really address the common case of a distro build which
is going to have Xen support enabled, and leaves the open question of
whether the native performance cost of pv-spinlocks is worth the
performance improvement on a loaded Xen system (10% saving of overall
system CPU when guests block rather than spin). Still it is a
reasonable short-term workaround.
[ Impact: fix pvops performance regression when running native ]
Analysed-by: "Xin Xiaohui" <xiaohui.xin@intel.com>
Analysed-by: "Li Xin" <xin.li@intel.com>
Analysed-by: "Nakajima Jun" <jun.nakajima@intel.com>
Signed-off-by: Jeremy Fitzhardinge <jeremy.fitzhardinge@citrix.com>
Acked-by: H. Peter Anvin <hpa@zytor.com>
Cc: Nick Piggin <npiggin@suse.de>
Cc: Xen-devel <xen-devel@lists.xensource.com>
LKML-Reference: <4A0B62F7.5030802@goop.org>
[ fixed the help text ]
Signed-off-by: Ingo Molnar <mingo@elte.hu>
2009-05-14 08:16:55 +08:00
|
|
|
#if defined(CONFIG_SMP) && defined(CONFIG_PARAVIRT_SPINLOCKS)
|
2008-07-09 20:33:33 +08:00
|
|
|
|
2015-04-25 02:56:38 +08:00
|
|
|
static __always_inline void pv_queued_spin_lock_slowpath(struct qspinlock *lock,
|
|
|
|
u32 val)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL2(lock.queued_spin_lock_slowpath, lock, val);
|
2015-04-25 02:56:38 +08:00
|
|
|
}
|
|
|
|
|
|
|
|
static __always_inline void pv_queued_spin_unlock(struct qspinlock *lock)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALLEE1(lock.queued_spin_unlock, lock);
|
2015-04-25 02:56:38 +08:00
|
|
|
}
|
|
|
|
|
|
|
|
static __always_inline void pv_wait(u8 *ptr, u8 val)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL2(lock.wait, ptr, val);
|
2015-04-25 02:56:38 +08:00
|
|
|
}
|
|
|
|
|
|
|
|
static __always_inline void pv_kick(int cpu)
|
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALL1(lock.kick, cpu);
|
2015-04-25 02:56:38 +08:00
|
|
|
}
|
|
|
|
|
2017-02-21 02:36:03 +08:00
|
|
|
static __always_inline bool pv_vcpu_is_preempted(long cpu)
|
2016-11-15 23:47:06 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
return PVOP_CALLEE1(bool, lock.vcpu_is_preempted, cpu);
|
2016-11-15 23:47:06 +08:00
|
|
|
}
|
|
|
|
|
2018-08-28 15:40:19 +08:00
|
|
|
void __raw_callee_save___native_queued_spin_unlock(struct qspinlock *lock);
|
|
|
|
bool __raw_callee_save___native_vcpu_is_preempted(long cpu);
|
|
|
|
|
2015-04-25 02:56:38 +08:00
|
|
|
#endif /* SMP && PARAVIRT_SPINLOCKS */
|
2008-07-09 20:33:33 +08:00
|
|
|
|
2008-01-30 20:32:07 +08:00
|
|
|
#ifdef CONFIG_X86_32
|
x86/paravirt: add register-saving thunks to reduce caller register pressure
Impact: Optimization
One of the problems with inserting a pile of C calls where previously
there were none is that the register pressure is greatly increased.
The C calling convention says that the caller must expect a certain
set of registers may be trashed by the callee, and that the callee can
use those registers without restriction. This includes the function
argument registers, and several others.
This patch seeks to alleviate this pressure by introducing wrapper
thunks that will do the register saving/restoring, so that the
callsite doesn't need to worry about it, but the callee function can
be conventional compiler-generated code. In many cases (particularly
performance-sensitive cases) the callee will be in assembler anyway,
and need not use the compiler's calling convention.
Standard calling convention is:
arguments return scratch
x86-32 eax edx ecx eax ?
x86-64 rdi rsi rdx rcx rax r8 r9 r10 r11
The thunk preserves all argument and scratch registers. The return
register is not preserved, and is available as a scratch register for
unwrapped callee code (and of course the return value).
Wrapped function pointers are themselves wrapped in a struct
paravirt_callee_save structure, in order to get some warning from the
compiler when functions with mismatched calling conventions are used.
The most common paravirt ops, both statically and dynamically, are
interrupt enable/disable/save/restore, so handle them first. This is
particularly easy since their calls are handled specially anyway.
XXX Deal with VMI. What's their calling convention?
Signed-off-by: H. Peter Anvin <hpa@zytor.com>
2009-01-29 06:35:05 +08:00
|
|
|
#define PV_SAVE_REGS "pushl %ecx; pushl %edx;"
|
|
|
|
#define PV_RESTORE_REGS "popl %edx; popl %ecx;"
|
|
|
|
|
|
|
|
/* save and restore all caller-save registers, except return value */
|
2009-01-31 15:17:23 +08:00
|
|
|
#define PV_SAVE_ALL_CALLER_REGS "pushl %ecx;"
|
|
|
|
#define PV_RESTORE_ALL_CALLER_REGS "popl %ecx;"
|
x86/paravirt: add register-saving thunks to reduce caller register pressure
Impact: Optimization
One of the problems with inserting a pile of C calls where previously
there were none is that the register pressure is greatly increased.
The C calling convention says that the caller must expect a certain
set of registers may be trashed by the callee, and that the callee can
use those registers without restriction. This includes the function
argument registers, and several others.
This patch seeks to alleviate this pressure by introducing wrapper
thunks that will do the register saving/restoring, so that the
callsite doesn't need to worry about it, but the callee function can
be conventional compiler-generated code. In many cases (particularly
performance-sensitive cases) the callee will be in assembler anyway,
and need not use the compiler's calling convention.
Standard calling convention is:
arguments return scratch
x86-32 eax edx ecx eax ?
x86-64 rdi rsi rdx rcx rax r8 r9 r10 r11
The thunk preserves all argument and scratch registers. The return
register is not preserved, and is available as a scratch register for
unwrapped callee code (and of course the return value).
Wrapped function pointers are themselves wrapped in a struct
paravirt_callee_save structure, in order to get some warning from the
compiler when functions with mismatched calling conventions are used.
The most common paravirt ops, both statically and dynamically, are
interrupt enable/disable/save/restore, so handle them first. This is
particularly easy since their calls are handled specially anyway.
XXX Deal with VMI. What's their calling convention?
Signed-off-by: H. Peter Anvin <hpa@zytor.com>
2009-01-29 06:35:05 +08:00
|
|
|
|
2008-01-30 20:32:07 +08:00
|
|
|
#define PV_FLAGS_ARG "0"
|
|
|
|
#define PV_EXTRA_CLOBBERS
|
|
|
|
#define PV_VEXTRA_CLOBBERS
|
|
|
|
#else
|
x86/paravirt: add register-saving thunks to reduce caller register pressure
Impact: Optimization
One of the problems with inserting a pile of C calls where previously
there were none is that the register pressure is greatly increased.
The C calling convention says that the caller must expect a certain
set of registers may be trashed by the callee, and that the callee can
use those registers without restriction. This includes the function
argument registers, and several others.
This patch seeks to alleviate this pressure by introducing wrapper
thunks that will do the register saving/restoring, so that the
callsite doesn't need to worry about it, but the callee function can
be conventional compiler-generated code. In many cases (particularly
performance-sensitive cases) the callee will be in assembler anyway,
and need not use the compiler's calling convention.
Standard calling convention is:
arguments return scratch
x86-32 eax edx ecx eax ?
x86-64 rdi rsi rdx rcx rax r8 r9 r10 r11
The thunk preserves all argument and scratch registers. The return
register is not preserved, and is available as a scratch register for
unwrapped callee code (and of course the return value).
Wrapped function pointers are themselves wrapped in a struct
paravirt_callee_save structure, in order to get some warning from the
compiler when functions with mismatched calling conventions are used.
The most common paravirt ops, both statically and dynamically, are
interrupt enable/disable/save/restore, so handle them first. This is
particularly easy since their calls are handled specially anyway.
XXX Deal with VMI. What's their calling convention?
Signed-off-by: H. Peter Anvin <hpa@zytor.com>
2009-01-29 06:35:05 +08:00
|
|
|
/* save and restore all caller-save registers, except return value */
|
|
|
|
#define PV_SAVE_ALL_CALLER_REGS \
|
|
|
|
"push %rcx;" \
|
|
|
|
"push %rdx;" \
|
|
|
|
"push %rsi;" \
|
|
|
|
"push %rdi;" \
|
|
|
|
"push %r8;" \
|
|
|
|
"push %r9;" \
|
|
|
|
"push %r10;" \
|
|
|
|
"push %r11;"
|
|
|
|
#define PV_RESTORE_ALL_CALLER_REGS \
|
|
|
|
"pop %r11;" \
|
|
|
|
"pop %r10;" \
|
|
|
|
"pop %r9;" \
|
|
|
|
"pop %r8;" \
|
|
|
|
"pop %rdi;" \
|
|
|
|
"pop %rsi;" \
|
|
|
|
"pop %rdx;" \
|
|
|
|
"pop %rcx;"
|
|
|
|
|
2008-01-30 20:32:07 +08:00
|
|
|
/* We save some registers, but all of them, that's too much. We clobber all
|
|
|
|
* caller saved registers but the argument parameter */
|
|
|
|
#define PV_SAVE_REGS "pushq %%rdi;"
|
|
|
|
#define PV_RESTORE_REGS "popq %%rdi;"
|
2008-07-09 06:07:12 +08:00
|
|
|
#define PV_EXTRA_CLOBBERS EXTRA_CLOBBERS, "rcx" , "rdx", "rsi"
|
|
|
|
#define PV_VEXTRA_CLOBBERS EXTRA_CLOBBERS, "rdi", "rcx" , "rdx", "rsi"
|
2008-01-30 20:32:07 +08:00
|
|
|
#define PV_FLAGS_ARG "D"
|
|
|
|
#endif
|
|
|
|
|
x86/paravirt: add register-saving thunks to reduce caller register pressure
Impact: Optimization
One of the problems with inserting a pile of C calls where previously
there were none is that the register pressure is greatly increased.
The C calling convention says that the caller must expect a certain
set of registers may be trashed by the callee, and that the callee can
use those registers without restriction. This includes the function
argument registers, and several others.
This patch seeks to alleviate this pressure by introducing wrapper
thunks that will do the register saving/restoring, so that the
callsite doesn't need to worry about it, but the callee function can
be conventional compiler-generated code. In many cases (particularly
performance-sensitive cases) the callee will be in assembler anyway,
and need not use the compiler's calling convention.
Standard calling convention is:
arguments return scratch
x86-32 eax edx ecx eax ?
x86-64 rdi rsi rdx rcx rax r8 r9 r10 r11
The thunk preserves all argument and scratch registers. The return
register is not preserved, and is available as a scratch register for
unwrapped callee code (and of course the return value).
Wrapped function pointers are themselves wrapped in a struct
paravirt_callee_save structure, in order to get some warning from the
compiler when functions with mismatched calling conventions are used.
The most common paravirt ops, both statically and dynamically, are
interrupt enable/disable/save/restore, so handle them first. This is
particularly easy since their calls are handled specially anyway.
XXX Deal with VMI. What's their calling convention?
Signed-off-by: H. Peter Anvin <hpa@zytor.com>
2009-01-29 06:35:05 +08:00
|
|
|
/*
|
|
|
|
* Generate a thunk around a function which saves all caller-save
|
|
|
|
* registers except for the return value. This allows C functions to
|
|
|
|
* be called from assembler code where fewer than normal registers are
|
|
|
|
* available. It may also help code generation around calls from C
|
|
|
|
* code if the common case doesn't use many registers.
|
|
|
|
*
|
|
|
|
* When a callee is wrapped in a thunk, the caller can assume that all
|
|
|
|
* arg regs and all scratch registers are preserved across the
|
|
|
|
* call. The return value in rax/eax will not be saved, even for void
|
|
|
|
* functions.
|
|
|
|
*/
|
2016-01-22 06:49:13 +08:00
|
|
|
#define PV_THUNK_NAME(func) "__raw_callee_save_" #func
|
x86/paravirt: add register-saving thunks to reduce caller register pressure
Impact: Optimization
One of the problems with inserting a pile of C calls where previously
there were none is that the register pressure is greatly increased.
The C calling convention says that the caller must expect a certain
set of registers may be trashed by the callee, and that the callee can
use those registers without restriction. This includes the function
argument registers, and several others.
This patch seeks to alleviate this pressure by introducing wrapper
thunks that will do the register saving/restoring, so that the
callsite doesn't need to worry about it, but the callee function can
be conventional compiler-generated code. In many cases (particularly
performance-sensitive cases) the callee will be in assembler anyway,
and need not use the compiler's calling convention.
Standard calling convention is:
arguments return scratch
x86-32 eax edx ecx eax ?
x86-64 rdi rsi rdx rcx rax r8 r9 r10 r11
The thunk preserves all argument and scratch registers. The return
register is not preserved, and is available as a scratch register for
unwrapped callee code (and of course the return value).
Wrapped function pointers are themselves wrapped in a struct
paravirt_callee_save structure, in order to get some warning from the
compiler when functions with mismatched calling conventions are used.
The most common paravirt ops, both statically and dynamically, are
interrupt enable/disable/save/restore, so handle them first. This is
particularly easy since their calls are handled specially anyway.
XXX Deal with VMI. What's their calling convention?
Signed-off-by: H. Peter Anvin <hpa@zytor.com>
2009-01-29 06:35:05 +08:00
|
|
|
#define PV_CALLEE_SAVE_REGS_THUNK(func) \
|
|
|
|
extern typeof(func) __raw_callee_save_##func; \
|
|
|
|
\
|
|
|
|
asm(".pushsection .text;" \
|
2016-01-22 06:49:13 +08:00
|
|
|
".globl " PV_THUNK_NAME(func) ";" \
|
|
|
|
".type " PV_THUNK_NAME(func) ", @function;" \
|
|
|
|
PV_THUNK_NAME(func) ":" \
|
|
|
|
FRAME_BEGIN \
|
x86/paravirt: add register-saving thunks to reduce caller register pressure
Impact: Optimization
One of the problems with inserting a pile of C calls where previously
there were none is that the register pressure is greatly increased.
The C calling convention says that the caller must expect a certain
set of registers may be trashed by the callee, and that the callee can
use those registers without restriction. This includes the function
argument registers, and several others.
This patch seeks to alleviate this pressure by introducing wrapper
thunks that will do the register saving/restoring, so that the
callsite doesn't need to worry about it, but the callee function can
be conventional compiler-generated code. In many cases (particularly
performance-sensitive cases) the callee will be in assembler anyway,
and need not use the compiler's calling convention.
Standard calling convention is:
arguments return scratch
x86-32 eax edx ecx eax ?
x86-64 rdi rsi rdx rcx rax r8 r9 r10 r11
The thunk preserves all argument and scratch registers. The return
register is not preserved, and is available as a scratch register for
unwrapped callee code (and of course the return value).
Wrapped function pointers are themselves wrapped in a struct
paravirt_callee_save structure, in order to get some warning from the
compiler when functions with mismatched calling conventions are used.
The most common paravirt ops, both statically and dynamically, are
interrupt enable/disable/save/restore, so handle them first. This is
particularly easy since their calls are handled specially anyway.
XXX Deal with VMI. What's their calling convention?
Signed-off-by: H. Peter Anvin <hpa@zytor.com>
2009-01-29 06:35:05 +08:00
|
|
|
PV_SAVE_ALL_CALLER_REGS \
|
|
|
|
"call " #func ";" \
|
|
|
|
PV_RESTORE_ALL_CALLER_REGS \
|
2016-01-22 06:49:13 +08:00
|
|
|
FRAME_END \
|
x86/paravirt: add register-saving thunks to reduce caller register pressure
Impact: Optimization
One of the problems with inserting a pile of C calls where previously
there were none is that the register pressure is greatly increased.
The C calling convention says that the caller must expect a certain
set of registers may be trashed by the callee, and that the callee can
use those registers without restriction. This includes the function
argument registers, and several others.
This patch seeks to alleviate this pressure by introducing wrapper
thunks that will do the register saving/restoring, so that the
callsite doesn't need to worry about it, but the callee function can
be conventional compiler-generated code. In many cases (particularly
performance-sensitive cases) the callee will be in assembler anyway,
and need not use the compiler's calling convention.
Standard calling convention is:
arguments return scratch
x86-32 eax edx ecx eax ?
x86-64 rdi rsi rdx rcx rax r8 r9 r10 r11
The thunk preserves all argument and scratch registers. The return
register is not preserved, and is available as a scratch register for
unwrapped callee code (and of course the return value).
Wrapped function pointers are themselves wrapped in a struct
paravirt_callee_save structure, in order to get some warning from the
compiler when functions with mismatched calling conventions are used.
The most common paravirt ops, both statically and dynamically, are
interrupt enable/disable/save/restore, so handle them first. This is
particularly easy since their calls are handled specially anyway.
XXX Deal with VMI. What's their calling convention?
Signed-off-by: H. Peter Anvin <hpa@zytor.com>
2009-01-29 06:35:05 +08:00
|
|
|
"ret;" \
|
|
|
|
".popsection")
|
|
|
|
|
|
|
|
/* Get a reference to a callee-save function */
|
|
|
|
#define PV_CALLEE_SAVE(func) \
|
|
|
|
((struct paravirt_callee_save) { __raw_callee_save_##func })
|
|
|
|
|
|
|
|
/* Promise that "func" already uses the right calling convention */
|
|
|
|
#define __PV_IS_CALLEE_SAVE(func) \
|
|
|
|
((struct paravirt_callee_save) { func })
|
|
|
|
|
2018-08-28 15:40:24 +08:00
|
|
|
#ifdef CONFIG_PARAVIRT_XXL
|
tracing: Force arch_local_irq_* notrace for paravirt
When running ktest.pl randconfig tests, I would sometimes trigger
a lockdep annotation bug (possible reason: unannotated irqs-on).
This triggering happened right after function tracer self test was
executed. After doing a config bisect I found that this was caused with
having function tracer, paravirt guest, prove locking, and rcu torture
all enabled.
The rcu torture just enhanced the likelyhood of triggering the bug.
Prove locking was needed, since it was the thing that was bugging.
Function tracer would trace and disable interrupts in all sorts
of funny places.
paravirt guest would turn arch_local_irq_* into functions that would
be traced.
Besides the fact that tracing arch_local_irq_* is just a bad idea,
this is what is happening.
The bug happened simply in the local_irq_restore() code:
if (raw_irqs_disabled_flags(flags)) { \
raw_local_irq_restore(flags); \
trace_hardirqs_off(); \
} else { \
trace_hardirqs_on(); \
raw_local_irq_restore(flags); \
} \
The raw_local_irq_restore() was defined as arch_local_irq_restore().
Now imagine, we are about to enable interrupts. We go into the else
case and call trace_hardirqs_on() which tells lockdep that we are enabling
interrupts, so it sets the current->hardirqs_enabled = 1.
Then we call raw_local_irq_restore() which calls arch_local_irq_restore()
which gets traced!
Now in the function tracer we disable interrupts with local_irq_save().
This is fine, but flags is stored that we have interrupts disabled.
When the function tracer calls local_irq_restore() it does it, but this
time with flags set as disabled, so we go into the if () path.
This keeps interrupts disabled and calls trace_hardirqs_off() which
sets current->hardirqs_enabled = 0.
When the tracer is finished and proceeds with the original code,
we enable interrupts but leave current->hardirqs_enabled as 0. Which
now breaks lockdeps internal processing.
Cc: Thomas Gleixner <tglx@linutronix.de>
Signed-off-by: Steven Rostedt <rostedt@goodmis.org>
2010-11-11 11:29:49 +08:00
|
|
|
static inline notrace unsigned long arch_local_save_flags(void)
|
2006-12-07 09:14:08 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
return PVOP_CALLEE0(unsigned long, irq.save_fl);
|
2006-12-07 09:14:08 +08:00
|
|
|
}
|
|
|
|
|
tracing: Force arch_local_irq_* notrace for paravirt
When running ktest.pl randconfig tests, I would sometimes trigger
a lockdep annotation bug (possible reason: unannotated irqs-on).
This triggering happened right after function tracer self test was
executed. After doing a config bisect I found that this was caused with
having function tracer, paravirt guest, prove locking, and rcu torture
all enabled.
The rcu torture just enhanced the likelyhood of triggering the bug.
Prove locking was needed, since it was the thing that was bugging.
Function tracer would trace and disable interrupts in all sorts
of funny places.
paravirt guest would turn arch_local_irq_* into functions that would
be traced.
Besides the fact that tracing arch_local_irq_* is just a bad idea,
this is what is happening.
The bug happened simply in the local_irq_restore() code:
if (raw_irqs_disabled_flags(flags)) { \
raw_local_irq_restore(flags); \
trace_hardirqs_off(); \
} else { \
trace_hardirqs_on(); \
raw_local_irq_restore(flags); \
} \
The raw_local_irq_restore() was defined as arch_local_irq_restore().
Now imagine, we are about to enable interrupts. We go into the else
case and call trace_hardirqs_on() which tells lockdep that we are enabling
interrupts, so it sets the current->hardirqs_enabled = 1.
Then we call raw_local_irq_restore() which calls arch_local_irq_restore()
which gets traced!
Now in the function tracer we disable interrupts with local_irq_save().
This is fine, but flags is stored that we have interrupts disabled.
When the function tracer calls local_irq_restore() it does it, but this
time with flags set as disabled, so we go into the if () path.
This keeps interrupts disabled and calls trace_hardirqs_off() which
sets current->hardirqs_enabled = 0.
When the tracer is finished and proceeds with the original code,
we enable interrupts but leave current->hardirqs_enabled as 0. Which
now breaks lockdeps internal processing.
Cc: Thomas Gleixner <tglx@linutronix.de>
Signed-off-by: Steven Rostedt <rostedt@goodmis.org>
2010-11-11 11:29:49 +08:00
|
|
|
static inline notrace void arch_local_irq_restore(unsigned long f)
|
2006-12-07 09:14:08 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALLEE1(irq.restore_fl, f);
|
2006-12-07 09:14:08 +08:00
|
|
|
}
|
|
|
|
|
tracing: Force arch_local_irq_* notrace for paravirt
When running ktest.pl randconfig tests, I would sometimes trigger
a lockdep annotation bug (possible reason: unannotated irqs-on).
This triggering happened right after function tracer self test was
executed. After doing a config bisect I found that this was caused with
having function tracer, paravirt guest, prove locking, and rcu torture
all enabled.
The rcu torture just enhanced the likelyhood of triggering the bug.
Prove locking was needed, since it was the thing that was bugging.
Function tracer would trace and disable interrupts in all sorts
of funny places.
paravirt guest would turn arch_local_irq_* into functions that would
be traced.
Besides the fact that tracing arch_local_irq_* is just a bad idea,
this is what is happening.
The bug happened simply in the local_irq_restore() code:
if (raw_irqs_disabled_flags(flags)) { \
raw_local_irq_restore(flags); \
trace_hardirqs_off(); \
} else { \
trace_hardirqs_on(); \
raw_local_irq_restore(flags); \
} \
The raw_local_irq_restore() was defined as arch_local_irq_restore().
Now imagine, we are about to enable interrupts. We go into the else
case and call trace_hardirqs_on() which tells lockdep that we are enabling
interrupts, so it sets the current->hardirqs_enabled = 1.
Then we call raw_local_irq_restore() which calls arch_local_irq_restore()
which gets traced!
Now in the function tracer we disable interrupts with local_irq_save().
This is fine, but flags is stored that we have interrupts disabled.
When the function tracer calls local_irq_restore() it does it, but this
time with flags set as disabled, so we go into the if () path.
This keeps interrupts disabled and calls trace_hardirqs_off() which
sets current->hardirqs_enabled = 0.
When the tracer is finished and proceeds with the original code,
we enable interrupts but leave current->hardirqs_enabled as 0. Which
now breaks lockdeps internal processing.
Cc: Thomas Gleixner <tglx@linutronix.de>
Signed-off-by: Steven Rostedt <rostedt@goodmis.org>
2010-11-11 11:29:49 +08:00
|
|
|
static inline notrace void arch_local_irq_disable(void)
|
2006-12-07 09:14:08 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALLEE0(irq.irq_disable);
|
2006-12-07 09:14:08 +08:00
|
|
|
}
|
|
|
|
|
tracing: Force arch_local_irq_* notrace for paravirt
When running ktest.pl randconfig tests, I would sometimes trigger
a lockdep annotation bug (possible reason: unannotated irqs-on).
This triggering happened right after function tracer self test was
executed. After doing a config bisect I found that this was caused with
having function tracer, paravirt guest, prove locking, and rcu torture
all enabled.
The rcu torture just enhanced the likelyhood of triggering the bug.
Prove locking was needed, since it was the thing that was bugging.
Function tracer would trace and disable interrupts in all sorts
of funny places.
paravirt guest would turn arch_local_irq_* into functions that would
be traced.
Besides the fact that tracing arch_local_irq_* is just a bad idea,
this is what is happening.
The bug happened simply in the local_irq_restore() code:
if (raw_irqs_disabled_flags(flags)) { \
raw_local_irq_restore(flags); \
trace_hardirqs_off(); \
} else { \
trace_hardirqs_on(); \
raw_local_irq_restore(flags); \
} \
The raw_local_irq_restore() was defined as arch_local_irq_restore().
Now imagine, we are about to enable interrupts. We go into the else
case and call trace_hardirqs_on() which tells lockdep that we are enabling
interrupts, so it sets the current->hardirqs_enabled = 1.
Then we call raw_local_irq_restore() which calls arch_local_irq_restore()
which gets traced!
Now in the function tracer we disable interrupts with local_irq_save().
This is fine, but flags is stored that we have interrupts disabled.
When the function tracer calls local_irq_restore() it does it, but this
time with flags set as disabled, so we go into the if () path.
This keeps interrupts disabled and calls trace_hardirqs_off() which
sets current->hardirqs_enabled = 0.
When the tracer is finished and proceeds with the original code,
we enable interrupts but leave current->hardirqs_enabled as 0. Which
now breaks lockdeps internal processing.
Cc: Thomas Gleixner <tglx@linutronix.de>
Signed-off-by: Steven Rostedt <rostedt@goodmis.org>
2010-11-11 11:29:49 +08:00
|
|
|
static inline notrace void arch_local_irq_enable(void)
|
2006-12-07 09:14:08 +08:00
|
|
|
{
|
2018-08-28 15:40:19 +08:00
|
|
|
PVOP_VCALLEE0(irq.irq_enable);
|
2006-12-07 09:14:08 +08:00
|
|
|
}
|
|
|
|
|
tracing: Force arch_local_irq_* notrace for paravirt
When running ktest.pl randconfig tests, I would sometimes trigger
a lockdep annotation bug (possible reason: unannotated irqs-on).
This triggering happened right after function tracer self test was
executed. After doing a config bisect I found that this was caused with
having function tracer, paravirt guest, prove locking, and rcu torture
all enabled.
The rcu torture just enhanced the likelyhood of triggering the bug.
Prove locking was needed, since it was the thing that was bugging.
Function tracer would trace and disable interrupts in all sorts
of funny places.
paravirt guest would turn arch_local_irq_* into functions that would
be traced.
Besides the fact that tracing arch_local_irq_* is just a bad idea,
this is what is happening.
The bug happened simply in the local_irq_restore() code:
if (raw_irqs_disabled_flags(flags)) { \
raw_local_irq_restore(flags); \
trace_hardirqs_off(); \
} else { \
trace_hardirqs_on(); \
raw_local_irq_restore(flags); \
} \
The raw_local_irq_restore() was defined as arch_local_irq_restore().
Now imagine, we are about to enable interrupts. We go into the else
case and call trace_hardirqs_on() which tells lockdep that we are enabling
interrupts, so it sets the current->hardirqs_enabled = 1.
Then we call raw_local_irq_restore() which calls arch_local_irq_restore()
which gets traced!
Now in the function tracer we disable interrupts with local_irq_save().
This is fine, but flags is stored that we have interrupts disabled.
When the function tracer calls local_irq_restore() it does it, but this
time with flags set as disabled, so we go into the if () path.
This keeps interrupts disabled and calls trace_hardirqs_off() which
sets current->hardirqs_enabled = 0.
When the tracer is finished and proceeds with the original code,
we enable interrupts but leave current->hardirqs_enabled as 0. Which
now breaks lockdeps internal processing.
Cc: Thomas Gleixner <tglx@linutronix.de>
Signed-off-by: Steven Rostedt <rostedt@goodmis.org>
2010-11-11 11:29:49 +08:00
|
|
|
static inline notrace unsigned long arch_local_irq_save(void)
|
2006-12-07 09:14:08 +08:00
|
|
|
{
|
|
|
|
unsigned long f;
|
|
|
|
|
2010-10-07 21:08:55 +08:00
|
|
|
f = arch_local_save_flags();
|
|
|
|
arch_local_irq_disable();
|
2006-12-07 09:14:08 +08:00
|
|
|
return f;
|
|
|
|
}
|
2018-08-28 15:40:24 +08:00
|
|
|
#endif
|
2006-12-07 09:14:08 +08:00
|
|
|
|
x86/paravirt: add hooks for spinlock operations
Ticket spinlocks have absolutely ghastly worst-case performance
characteristics in a virtual environment. If there is any contention
for physical CPUs (ie, there are more runnable vcpus than cpus), then
ticket locks can cause the system to end up spending 90+% of its time
spinning.
The problem is that (v)cpus waiting on a ticket spinlock will be
granted access to the lock in strict order they got their tickets. If
the hypervisor scheduler doesn't give the vcpus time in that order,
they will burn timeslices waiting for the scheduler to give the right
vcpu some time. In the worst case it could take O(n^2) vcpu scheduler
timeslices for everyone waiting on the lock to get it, not counting
new cpus trying to take the lock while the log-jam is sorted out.
These hooks allow a paravirt backend to replace the spinlock
implementation.
At the very least, this could revert the implementation back to the
old lock algorithm, which allows the next scheduled vcpu to take the
lock, and has basically fairly good performance.
It also allows the spinlocks to take advantages of the hypervisor
features to make locks more efficient (spin and block, for example).
The cost to native execution is an extra direct call when using a
spinlock function. There's no overhead if CONFIG_PARAVIRT is turned
off.
The lock structure is fixed at a single "unsigned int", initialized to
zero, but the spinlock implementation can use it as it wishes.
Thanks to Thomas Friebel's Xen Summit talk "Preventing Guests from
Spinning Around" for pointing out this problem.
Signed-off-by: Jeremy Fitzhardinge <jeremy.fitzhardinge@citrix.com>
Cc: Jens Axboe <axboe@kernel.dk>
Cc: Peter Zijlstra <a.p.zijlstra@chello.nl>
Cc: Christoph Lameter <clameter@linux-foundation.org>
Cc: Petr Tesarik <ptesarik@suse.cz>
Cc: Virtualization <virtualization@lists.linux-foundation.org>
Cc: Xen devel <xen-devel@lists.xensource.com>
Cc: Thomas Friebel <thomas.friebel@amd.com>
Cc: Nick Piggin <nickpiggin@yahoo.com.au>
Signed-off-by: Ingo Molnar <mingo@elte.hu>
2008-07-08 03:07:50 +08:00
|
|
|
|
2007-05-03 01:27:14 +08:00
|
|
|
/* Make sure as little as possible of this mess escapes. */
|
2007-05-03 01:27:14 +08:00
|
|
|
#undef PARAVIRT_CALL
|
2007-05-03 01:27:15 +08:00
|
|
|
#undef __PVOP_CALL
|
|
|
|
#undef __PVOP_VCALL
|
2007-05-03 01:27:14 +08:00
|
|
|
#undef PVOP_VCALL0
|
|
|
|
#undef PVOP_CALL0
|
|
|
|
#undef PVOP_VCALL1
|
|
|
|
#undef PVOP_CALL1
|
|
|
|
#undef PVOP_VCALL2
|
|
|
|
#undef PVOP_CALL2
|
|
|
|
#undef PVOP_VCALL3
|
|
|
|
#undef PVOP_CALL3
|
|
|
|
#undef PVOP_VCALL4
|
|
|
|
#undef PVOP_CALL4
|
2006-12-07 09:14:08 +08:00
|
|
|
|
2009-08-20 19:19:57 +08:00
|
|
|
extern void default_banner(void);
|
|
|
|
|
2006-12-07 09:14:07 +08:00
|
|
|
#else /* __ASSEMBLY__ */
|
|
|
|
|
2018-08-28 15:40:18 +08:00
|
|
|
#define _PVSITE(ptype, ops, word, algn) \
|
2006-12-07 09:14:08 +08:00
|
|
|
771:; \
|
|
|
|
ops; \
|
|
|
|
772:; \
|
|
|
|
.pushsection .parainstructions,"a"; \
|
2008-01-30 20:32:06 +08:00
|
|
|
.align algn; \
|
|
|
|
word 771b; \
|
2006-12-07 09:14:08 +08:00
|
|
|
.byte ptype; \
|
|
|
|
.byte 772b-771b; \
|
|
|
|
.popsection
|
|
|
|
|
2008-01-30 20:32:06 +08:00
|
|
|
|
2009-01-29 06:35:04 +08:00
|
|
|
#define COND_PUSH(set, mask, reg) \
|
x86/paravirt: add register-saving thunks to reduce caller register pressure
Impact: Optimization
One of the problems with inserting a pile of C calls where previously
there were none is that the register pressure is greatly increased.
The C calling convention says that the caller must expect a certain
set of registers may be trashed by the callee, and that the callee can
use those registers without restriction. This includes the function
argument registers, and several others.
This patch seeks to alleviate this pressure by introducing wrapper
thunks that will do the register saving/restoring, so that the
callsite doesn't need to worry about it, but the callee function can
be conventional compiler-generated code. In many cases (particularly
performance-sensitive cases) the callee will be in assembler anyway,
and need not use the compiler's calling convention.
Standard calling convention is:
arguments return scratch
x86-32 eax edx ecx eax ?
x86-64 rdi rsi rdx rcx rax r8 r9 r10 r11
The thunk preserves all argument and scratch registers. The return
register is not preserved, and is available as a scratch register for
unwrapped callee code (and of course the return value).
Wrapped function pointers are themselves wrapped in a struct
paravirt_callee_save structure, in order to get some warning from the
compiler when functions with mismatched calling conventions are used.
The most common paravirt ops, both statically and dynamically, are
interrupt enable/disable/save/restore, so handle them first. This is
particularly easy since their calls are handled specially anyway.
XXX Deal with VMI. What's their calling convention?
Signed-off-by: H. Peter Anvin <hpa@zytor.com>
2009-01-29 06:35:05 +08:00
|
|
|
.if ((~(set)) & mask); push %reg; .endif
|
2009-01-29 06:35:04 +08:00
|
|
|
#define COND_POP(set, mask, reg) \
|
x86/paravirt: add register-saving thunks to reduce caller register pressure
Impact: Optimization
One of the problems with inserting a pile of C calls where previously
there were none is that the register pressure is greatly increased.
The C calling convention says that the caller must expect a certain
set of registers may be trashed by the callee, and that the callee can
use those registers without restriction. This includes the function
argument registers, and several others.
This patch seeks to alleviate this pressure by introducing wrapper
thunks that will do the register saving/restoring, so that the
callsite doesn't need to worry about it, but the callee function can
be conventional compiler-generated code. In many cases (particularly
performance-sensitive cases) the callee will be in assembler anyway,
and need not use the compiler's calling convention.
Standard calling convention is:
arguments return scratch
x86-32 eax edx ecx eax ?
x86-64 rdi rsi rdx rcx rax r8 r9 r10 r11
The thunk preserves all argument and scratch registers. The return
register is not preserved, and is available as a scratch register for
unwrapped callee code (and of course the return value).
Wrapped function pointers are themselves wrapped in a struct
paravirt_callee_save structure, in order to get some warning from the
compiler when functions with mismatched calling conventions are used.
The most common paravirt ops, both statically and dynamically, are
interrupt enable/disable/save/restore, so handle them first. This is
particularly easy since their calls are handled specially anyway.
XXX Deal with VMI. What's their calling convention?
Signed-off-by: H. Peter Anvin <hpa@zytor.com>
2009-01-29 06:35:05 +08:00
|
|
|
.if ((~(set)) & mask); pop %reg; .endif
|
2009-01-29 06:35:04 +08:00
|
|
|
|
2008-01-30 20:32:06 +08:00
|
|
|
#ifdef CONFIG_X86_64
|
2009-01-29 06:35:04 +08:00
|
|
|
|
|
|
|
#define PV_SAVE_REGS(set) \
|
|
|
|
COND_PUSH(set, CLBR_RAX, rax); \
|
|
|
|
COND_PUSH(set, CLBR_RCX, rcx); \
|
|
|
|
COND_PUSH(set, CLBR_RDX, rdx); \
|
|
|
|
COND_PUSH(set, CLBR_RSI, rsi); \
|
|
|
|
COND_PUSH(set, CLBR_RDI, rdi); \
|
|
|
|
COND_PUSH(set, CLBR_R8, r8); \
|
|
|
|
COND_PUSH(set, CLBR_R9, r9); \
|
|
|
|
COND_PUSH(set, CLBR_R10, r10); \
|
|
|
|
COND_PUSH(set, CLBR_R11, r11)
|
|
|
|
#define PV_RESTORE_REGS(set) \
|
|
|
|
COND_POP(set, CLBR_R11, r11); \
|
|
|
|
COND_POP(set, CLBR_R10, r10); \
|
|
|
|
COND_POP(set, CLBR_R9, r9); \
|
|
|
|
COND_POP(set, CLBR_R8, r8); \
|
|
|
|
COND_POP(set, CLBR_RDI, rdi); \
|
|
|
|
COND_POP(set, CLBR_RSI, rsi); \
|
|
|
|
COND_POP(set, CLBR_RDX, rdx); \
|
|
|
|
COND_POP(set, CLBR_RCX, rcx); \
|
|
|
|
COND_POP(set, CLBR_RAX, rax)
|
|
|
|
|
2018-08-28 15:40:19 +08:00
|
|
|
#define PARA_PATCH(off) ((off) / 8)
|
2018-08-28 15:40:18 +08:00
|
|
|
#define PARA_SITE(ptype, ops) _PVSITE(ptype, ops, .quad, 8)
|
2008-06-25 12:19:15 +08:00
|
|
|
#define PARA_INDIRECT(addr) *addr(%rip)
|
2008-01-30 20:32:06 +08:00
|
|
|
#else
|
2009-01-29 06:35:04 +08:00
|
|
|
#define PV_SAVE_REGS(set) \
|
|
|
|
COND_PUSH(set, CLBR_EAX, eax); \
|
|
|
|
COND_PUSH(set, CLBR_EDI, edi); \
|
|
|
|
COND_PUSH(set, CLBR_ECX, ecx); \
|
|
|
|
COND_PUSH(set, CLBR_EDX, edx)
|
|
|
|
#define PV_RESTORE_REGS(set) \
|
|
|
|
COND_POP(set, CLBR_EDX, edx); \
|
|
|
|
COND_POP(set, CLBR_ECX, ecx); \
|
|
|
|
COND_POP(set, CLBR_EDI, edi); \
|
|
|
|
COND_POP(set, CLBR_EAX, eax)
|
|
|
|
|
2018-08-28 15:40:19 +08:00
|
|
|
#define PARA_PATCH(off) ((off) / 4)
|
2018-08-28 15:40:18 +08:00
|
|
|
#define PARA_SITE(ptype, ops) _PVSITE(ptype, ops, .long, 4)
|
2008-06-25 12:19:15 +08:00
|
|
|
#define PARA_INDIRECT(addr) *%cs:addr
|
2008-01-30 20:32:06 +08:00
|
|
|
#endif
|
|
|
|
|
2018-08-28 15:40:23 +08:00
|
|
|
#ifdef CONFIG_PARAVIRT_XXL
|
2007-10-17 02:51:29 +08:00
|
|
|
#define INTERRUPT_RETURN \
|
2018-08-28 15:40:19 +08:00
|
|
|
PARA_SITE(PARA_PATCH(PV_CPU_iret), \
|
2018-08-28 15:40:18 +08:00
|
|
|
ANNOTATE_RETPOLINE_SAFE; \
|
2018-08-28 15:40:19 +08:00
|
|
|
jmp PARA_INDIRECT(pv_ops+PV_CPU_iret);)
|
2007-05-03 01:27:14 +08:00
|
|
|
|
|
|
|
#define DISABLE_INTERRUPTS(clobbers) \
|
2018-08-28 15:40:19 +08:00
|
|
|
PARA_SITE(PARA_PATCH(PV_IRQ_irq_disable), \
|
x86/paravirt: add register-saving thunks to reduce caller register pressure
Impact: Optimization
One of the problems with inserting a pile of C calls where previously
there were none is that the register pressure is greatly increased.
The C calling convention says that the caller must expect a certain
set of registers may be trashed by the callee, and that the callee can
use those registers without restriction. This includes the function
argument registers, and several others.
This patch seeks to alleviate this pressure by introducing wrapper
thunks that will do the register saving/restoring, so that the
callsite doesn't need to worry about it, but the callee function can
be conventional compiler-generated code. In many cases (particularly
performance-sensitive cases) the callee will be in assembler anyway,
and need not use the compiler's calling convention.
Standard calling convention is:
arguments return scratch
x86-32 eax edx ecx eax ?
x86-64 rdi rsi rdx rcx rax r8 r9 r10 r11
The thunk preserves all argument and scratch registers. The return
register is not preserved, and is available as a scratch register for
unwrapped callee code (and of course the return value).
Wrapped function pointers are themselves wrapped in a struct
paravirt_callee_save structure, in order to get some warning from the
compiler when functions with mismatched calling conventions are used.
The most common paravirt ops, both statically and dynamically, are
interrupt enable/disable/save/restore, so handle them first. This is
particularly easy since their calls are handled specially anyway.
XXX Deal with VMI. What's their calling convention?
Signed-off-by: H. Peter Anvin <hpa@zytor.com>
2009-01-29 06:35:05 +08:00
|
|
|
PV_SAVE_REGS(clobbers | CLBR_CALLEE_SAVE); \
|
2018-08-28 15:40:18 +08:00
|
|
|
ANNOTATE_RETPOLINE_SAFE; \
|
2018-08-28 15:40:19 +08:00
|
|
|
call PARA_INDIRECT(pv_ops+PV_IRQ_irq_disable); \
|
x86/paravirt: add register-saving thunks to reduce caller register pressure
Impact: Optimization
One of the problems with inserting a pile of C calls where previously
there were none is that the register pressure is greatly increased.
The C calling convention says that the caller must expect a certain
set of registers may be trashed by the callee, and that the callee can
use those registers without restriction. This includes the function
argument registers, and several others.
This patch seeks to alleviate this pressure by introducing wrapper
thunks that will do the register saving/restoring, so that the
callsite doesn't need to worry about it, but the callee function can
be conventional compiler-generated code. In many cases (particularly
performance-sensitive cases) the callee will be in assembler anyway,
and need not use the compiler's calling convention.
Standard calling convention is:
arguments return scratch
x86-32 eax edx ecx eax ?
x86-64 rdi rsi rdx rcx rax r8 r9 r10 r11
The thunk preserves all argument and scratch registers. The return
register is not preserved, and is available as a scratch register for
unwrapped callee code (and of course the return value).
Wrapped function pointers are themselves wrapped in a struct
paravirt_callee_save structure, in order to get some warning from the
compiler when functions with mismatched calling conventions are used.
The most common paravirt ops, both statically and dynamically, are
interrupt enable/disable/save/restore, so handle them first. This is
particularly easy since their calls are handled specially anyway.
XXX Deal with VMI. What's their calling convention?
Signed-off-by: H. Peter Anvin <hpa@zytor.com>
2009-01-29 06:35:05 +08:00
|
|
|
PV_RESTORE_REGS(clobbers | CLBR_CALLEE_SAVE);)
|
2007-05-03 01:27:14 +08:00
|
|
|
|
|
|
|
#define ENABLE_INTERRUPTS(clobbers) \
|
2018-08-28 15:40:19 +08:00
|
|
|
PARA_SITE(PARA_PATCH(PV_IRQ_irq_enable), \
|
x86/paravirt: add register-saving thunks to reduce caller register pressure
Impact: Optimization
One of the problems with inserting a pile of C calls where previously
there were none is that the register pressure is greatly increased.
The C calling convention says that the caller must expect a certain
set of registers may be trashed by the callee, and that the callee can
use those registers without restriction. This includes the function
argument registers, and several others.
This patch seeks to alleviate this pressure by introducing wrapper
thunks that will do the register saving/restoring, so that the
callsite doesn't need to worry about it, but the callee function can
be conventional compiler-generated code. In many cases (particularly
performance-sensitive cases) the callee will be in assembler anyway,
and need not use the compiler's calling convention.
Standard calling convention is:
arguments return scratch
x86-32 eax edx ecx eax ?
x86-64 rdi rsi rdx rcx rax r8 r9 r10 r11
The thunk preserves all argument and scratch registers. The return
register is not preserved, and is available as a scratch register for
unwrapped callee code (and of course the return value).
Wrapped function pointers are themselves wrapped in a struct
paravirt_callee_save structure, in order to get some warning from the
compiler when functions with mismatched calling conventions are used.
The most common paravirt ops, both statically and dynamically, are
interrupt enable/disable/save/restore, so handle them first. This is
particularly easy since their calls are handled specially anyway.
XXX Deal with VMI. What's their calling convention?
Signed-off-by: H. Peter Anvin <hpa@zytor.com>
2009-01-29 06:35:05 +08:00
|
|
|
PV_SAVE_REGS(clobbers | CLBR_CALLEE_SAVE); \
|
2018-08-28 15:40:18 +08:00
|
|
|
ANNOTATE_RETPOLINE_SAFE; \
|
2018-08-28 15:40:19 +08:00
|
|
|
call PARA_INDIRECT(pv_ops+PV_IRQ_irq_enable); \
|
x86/paravirt: add register-saving thunks to reduce caller register pressure
Impact: Optimization
One of the problems with inserting a pile of C calls where previously
there were none is that the register pressure is greatly increased.
The C calling convention says that the caller must expect a certain
set of registers may be trashed by the callee, and that the callee can
use those registers without restriction. This includes the function
argument registers, and several others.
This patch seeks to alleviate this pressure by introducing wrapper
thunks that will do the register saving/restoring, so that the
callsite doesn't need to worry about it, but the callee function can
be conventional compiler-generated code. In many cases (particularly
performance-sensitive cases) the callee will be in assembler anyway,
and need not use the compiler's calling convention.
Standard calling convention is:
arguments return scratch
x86-32 eax edx ecx eax ?
x86-64 rdi rsi rdx rcx rax r8 r9 r10 r11
The thunk preserves all argument and scratch registers. The return
register is not preserved, and is available as a scratch register for
unwrapped callee code (and of course the return value).
Wrapped function pointers are themselves wrapped in a struct
paravirt_callee_save structure, in order to get some warning from the
compiler when functions with mismatched calling conventions are used.
The most common paravirt ops, both statically and dynamically, are
interrupt enable/disable/save/restore, so handle them first. This is
particularly easy since their calls are handled specially anyway.
XXX Deal with VMI. What's their calling convention?
Signed-off-by: H. Peter Anvin <hpa@zytor.com>
2009-01-29 06:35:05 +08:00
|
|
|
PV_RESTORE_REGS(clobbers | CLBR_CALLEE_SAVE);)
|
2018-08-28 15:40:24 +08:00
|
|
|
#endif
|
2007-05-03 01:27:14 +08:00
|
|
|
|
2018-08-28 15:40:20 +08:00
|
|
|
#ifdef CONFIG_X86_64
|
2018-08-28 15:40:23 +08:00
|
|
|
#ifdef CONFIG_PARAVIRT_XXL
|
2008-06-25 12:19:30 +08:00
|
|
|
/*
|
|
|
|
* If swapgs is used while the userspace stack is still current,
|
|
|
|
* there's no way to call a pvop. The PV replacement *must* be
|
|
|
|
* inlined, or the swapgs instruction must be trapped and emulated.
|
|
|
|
*/
|
|
|
|
#define SWAPGS_UNSAFE_STACK \
|
2018-08-28 15:40:19 +08:00
|
|
|
PARA_SITE(PARA_PATCH(PV_CPU_swapgs), swapgs)
|
2008-06-25 12:19:30 +08:00
|
|
|
|
2009-01-29 06:35:04 +08:00
|
|
|
/*
|
|
|
|
* Note: swapgs is very special, and in practise is either going to be
|
|
|
|
* implemented with a single "swapgs" instruction or something very
|
|
|
|
* special. Either way, we don't need to save any registers for
|
|
|
|
* it.
|
|
|
|
*/
|
2008-01-30 20:32:08 +08:00
|
|
|
#define SWAPGS \
|
2018-08-28 15:40:19 +08:00
|
|
|
PARA_SITE(PARA_PATCH(PV_CPU_swapgs), \
|
2018-08-28 15:40:18 +08:00
|
|
|
ANNOTATE_RETPOLINE_SAFE; \
|
2018-08-28 15:40:19 +08:00
|
|
|
call PARA_INDIRECT(pv_ops+PV_CPU_swapgs); \
|
2008-01-30 20:32:08 +08:00
|
|
|
)
|
2018-08-28 15:40:23 +08:00
|
|
|
#endif
|
2008-01-30 20:32:08 +08:00
|
|
|
|
2012-04-19 08:16:48 +08:00
|
|
|
#define GET_CR2_INTO_RAX \
|
2018-01-17 23:58:11 +08:00
|
|
|
ANNOTATE_RETPOLINE_SAFE; \
|
2018-08-28 15:40:19 +08:00
|
|
|
call PARA_INDIRECT(pv_ops+PV_MMU_read_cr2);
|
2008-01-30 20:32:07 +08:00
|
|
|
|
2018-08-28 15:40:23 +08:00
|
|
|
#ifdef CONFIG_PARAVIRT_XXL
|
2008-06-25 12:19:28 +08:00
|
|
|
#define USERGS_SYSRET64 \
|
2018-08-28 15:40:19 +08:00
|
|
|
PARA_SITE(PARA_PATCH(PV_CPU_usergs_sysret64), \
|
2018-08-28 15:40:18 +08:00
|
|
|
ANNOTATE_RETPOLINE_SAFE; \
|
2018-08-28 15:40:19 +08:00
|
|
|
jmp PARA_INDIRECT(pv_ops+PV_CPU_usergs_sysret64);)
|
2017-12-04 22:07:07 +08:00
|
|
|
|
|
|
|
#ifdef CONFIG_DEBUG_ENTRY
|
|
|
|
#define SAVE_FLAGS(clobbers) \
|
2018-08-28 15:40:19 +08:00
|
|
|
PARA_SITE(PARA_PATCH(PV_IRQ_save_fl), \
|
2017-12-04 22:07:07 +08:00
|
|
|
PV_SAVE_REGS(clobbers | CLBR_CALLEE_SAVE); \
|
2018-08-28 15:40:18 +08:00
|
|
|
ANNOTATE_RETPOLINE_SAFE; \
|
2018-08-28 15:40:19 +08:00
|
|
|
call PARA_INDIRECT(pv_ops+PV_IRQ_save_fl); \
|
2017-12-04 22:07:07 +08:00
|
|
|
PV_RESTORE_REGS(clobbers | CLBR_CALLEE_SAVE);)
|
|
|
|
#endif
|
2018-09-05 13:37:20 +08:00
|
|
|
#endif
|
2017-12-04 22:07:07 +08:00
|
|
|
|
2008-06-25 12:19:28 +08:00
|
|
|
#endif /* CONFIG_X86_32 */
|
2006-12-07 09:14:08 +08:00
|
|
|
|
2006-12-07 09:14:07 +08:00
|
|
|
#endif /* __ASSEMBLY__ */
|
2009-08-20 19:19:57 +08:00
|
|
|
#else /* CONFIG_PARAVIRT */
|
|
|
|
# define default_banner x86_init_noop
|
2018-08-28 15:40:25 +08:00
|
|
|
#endif /* !CONFIG_PARAVIRT */
|
|
|
|
|
2014-11-19 02:23:49 +08:00
|
|
|
#ifndef __ASSEMBLY__
|
2018-08-28 15:40:25 +08:00
|
|
|
#ifndef CONFIG_PARAVIRT_XXL
|
2014-11-19 02:23:49 +08:00
|
|
|
static inline void paravirt_arch_dup_mmap(struct mm_struct *oldmm,
|
|
|
|
struct mm_struct *mm)
|
|
|
|
{
|
|
|
|
}
|
2018-08-28 15:40:25 +08:00
|
|
|
#endif
|
2014-11-19 02:23:49 +08:00
|
|
|
|
2018-08-28 15:40:25 +08:00
|
|
|
#ifndef CONFIG_PARAVIRT
|
2014-11-19 02:23:49 +08:00
|
|
|
static inline void paravirt_arch_exit_mmap(struct mm_struct *mm)
|
|
|
|
{
|
|
|
|
}
|
2018-08-28 15:40:25 +08:00
|
|
|
#endif
|
2014-11-19 02:23:49 +08:00
|
|
|
#endif /* __ASSEMBLY__ */
|
2008-10-23 13:26:29 +08:00
|
|
|
#endif /* _ASM_X86_PARAVIRT_H */
|