2012-10-05 08:16:38 +08:00
|
|
|
/* Copyright (c) 2012 Coraid, Inc. See COPYING for GPL terms. */
|
2005-04-17 06:20:36 +08:00
|
|
|
/*
|
|
|
|
* aoecmd.c
|
|
|
|
* Filesystem request handling methods
|
|
|
|
*/
|
|
|
|
|
2009-04-02 03:42:24 +08:00
|
|
|
#include <linux/ata.h>
|
include cleanup: Update gfp.h and slab.h includes to prepare for breaking implicit slab.h inclusion from percpu.h
percpu.h is included by sched.h and module.h and thus ends up being
included when building most .c files. percpu.h includes slab.h which
in turn includes gfp.h making everything defined by the two files
universally available and complicating inclusion dependencies.
percpu.h -> slab.h dependency is about to be removed. Prepare for
this change by updating users of gfp and slab facilities include those
headers directly instead of assuming availability. As this conversion
needs to touch large number of source files, the following script is
used as the basis of conversion.
http://userweb.kernel.org/~tj/misc/slabh-sweep.py
The script does the followings.
* Scan files for gfp and slab usages and update includes such that
only the necessary includes are there. ie. if only gfp is used,
gfp.h, if slab is used, slab.h.
* When the script inserts a new include, it looks at the include
blocks and try to put the new include such that its order conforms
to its surrounding. It's put in the include block which contains
core kernel includes, in the same order that the rest are ordered -
alphabetical, Christmas tree, rev-Xmas-tree or at the end if there
doesn't seem to be any matching order.
* If the script can't find a place to put a new include (mostly
because the file doesn't have fitting include block), it prints out
an error message indicating which .h file needs to be added to the
file.
The conversion was done in the following steps.
1. The initial automatic conversion of all .c files updated slightly
over 4000 files, deleting around 700 includes and adding ~480 gfp.h
and ~3000 slab.h inclusions. The script emitted errors for ~400
files.
2. Each error was manually checked. Some didn't need the inclusion,
some needed manual addition while adding it to implementation .h or
embedding .c file was more appropriate for others. This step added
inclusions to around 150 files.
3. The script was run again and the output was compared to the edits
from #2 to make sure no file was left behind.
4. Several build tests were done and a couple of problems were fixed.
e.g. lib/decompress_*.c used malloc/free() wrappers around slab
APIs requiring slab.h to be added manually.
5. The script was run on all .h files but without automatically
editing them as sprinkling gfp.h and slab.h inclusions around .h
files could easily lead to inclusion dependency hell. Most gfp.h
inclusion directives were ignored as stuff from gfp.h was usually
wildly available and often used in preprocessor macros. Each
slab.h inclusion directive was examined and added manually as
necessary.
6. percpu.h was updated not to include slab.h.
7. Build test were done on the following configurations and failures
were fixed. CONFIG_GCOV_KERNEL was turned off for all tests (as my
distributed build env didn't work with gcov compiles) and a few
more options had to be turned off depending on archs to make things
build (like ipr on powerpc/64 which failed due to missing writeq).
* x86 and x86_64 UP and SMP allmodconfig and a custom test config.
* powerpc and powerpc64 SMP allmodconfig
* sparc and sparc64 SMP allmodconfig
* ia64 SMP allmodconfig
* s390 SMP allmodconfig
* alpha SMP allmodconfig
* um on x86_64 SMP allmodconfig
8. percpu.h modifications were reverted so that it could be applied as
a separate patch and serve as bisection point.
Given the fact that I had only a couple of failures from tests on step
6, I'm fairly confident about the coverage of this conversion patch.
If there is a breakage, it's likely to be something in one of the arch
headers which should be easily discoverable easily on most builds of
the specific arch.
Signed-off-by: Tejun Heo <tj@kernel.org>
Guess-its-ok-by: Christoph Lameter <cl@linux-foundation.org>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Lee Schermerhorn <Lee.Schermerhorn@hp.com>
2010-03-24 16:04:11 +08:00
|
|
|
#include <linux/slab.h>
|
2005-04-17 06:20:36 +08:00
|
|
|
#include <linux/hdreg.h>
|
|
|
|
#include <linux/blkdev.h>
|
|
|
|
#include <linux/skbuff.h>
|
|
|
|
#include <linux/netdevice.h>
|
2006-01-20 02:46:19 +08:00
|
|
|
#include <linux/genhd.h>
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
#include <linux/moduleparam.h>
|
2012-10-05 08:16:21 +08:00
|
|
|
#include <linux/workqueue.h>
|
|
|
|
#include <linux/kthread.h>
|
2007-09-18 02:56:21 +08:00
|
|
|
#include <net/net_namespace.h>
|
2005-09-30 00:47:40 +08:00
|
|
|
#include <asm/unaligned.h>
|
2012-10-05 08:16:21 +08:00
|
|
|
#include <linux/uio.h>
|
2005-04-17 06:20:36 +08:00
|
|
|
#include "aoe.h"
|
|
|
|
|
2012-10-05 08:16:21 +08:00
|
|
|
#define MAXIOC (8192) /* default meant to avoid most soft lockups */
|
|
|
|
|
|
|
|
static void ktcomplete(struct frame *, struct sk_buff *);
|
|
|
|
|
2012-10-05 08:16:23 +08:00
|
|
|
static struct buf *nextbuf(struct aoedev *);
|
|
|
|
|
2006-09-21 02:36:50 +08:00
|
|
|
static int aoe_deadsecs = 60 * 3;
|
|
|
|
module_param(aoe_deadsecs, int, 0644);
|
|
|
|
MODULE_PARM_DESC(aoe_deadsecs, "After aoe_deadsecs seconds, give up and fail dev.");
|
2005-04-17 06:20:36 +08:00
|
|
|
|
2008-02-08 20:20:07 +08:00
|
|
|
static int aoe_maxout = 16;
|
|
|
|
module_param(aoe_maxout, int, 0644);
|
|
|
|
MODULE_PARM_DESC(aoe_maxout,
|
|
|
|
"Only aoe_maxout outstanding packets for every MAC on eX.Y.");
|
|
|
|
|
2012-10-05 08:16:21 +08:00
|
|
|
static wait_queue_head_t ktiowq;
|
|
|
|
static struct ktstate kts;
|
|
|
|
|
|
|
|
/* io completion queue */
|
|
|
|
static struct {
|
|
|
|
struct list_head head;
|
|
|
|
spinlock_t lock;
|
|
|
|
} iocq;
|
|
|
|
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
static struct sk_buff *
|
2006-09-21 02:36:49 +08:00
|
|
|
new_skb(ulong len)
|
2005-04-17 06:20:36 +08:00
|
|
|
{
|
|
|
|
struct sk_buff *skb;
|
|
|
|
|
|
|
|
skb = alloc_skb(len, GFP_ATOMIC);
|
|
|
|
if (skb) {
|
2007-03-20 06:30:44 +08:00
|
|
|
skb_reset_mac_header(skb);
|
2007-04-11 11:45:18 +08:00
|
|
|
skb_reset_network_header(skb);
|
2005-04-17 06:20:36 +08:00
|
|
|
skb->protocol = __constant_htons(ETH_P_AOE);
|
2012-09-19 23:46:39 +08:00
|
|
|
skb_checksum_none_assert(skb);
|
2005-04-17 06:20:36 +08:00
|
|
|
}
|
|
|
|
return skb;
|
|
|
|
}
|
|
|
|
|
|
|
|
static struct frame *
|
2012-10-05 08:16:33 +08:00
|
|
|
getframe(struct aoedev *d, u32 tag)
|
2005-04-17 06:20:36 +08:00
|
|
|
{
|
2012-10-05 08:16:21 +08:00
|
|
|
struct frame *f;
|
|
|
|
struct list_head *head, *pos, *nx;
|
|
|
|
u32 n;
|
2005-04-17 06:20:36 +08:00
|
|
|
|
2012-10-05 08:16:21 +08:00
|
|
|
n = tag % NFACTIVE;
|
2012-10-05 08:16:33 +08:00
|
|
|
head = &d->factive[n];
|
2012-10-05 08:16:21 +08:00
|
|
|
list_for_each_safe(pos, nx, head) {
|
|
|
|
f = list_entry(pos, struct frame, head);
|
|
|
|
if (f->tag == tag) {
|
|
|
|
list_del(pos);
|
2005-04-17 06:20:36 +08:00
|
|
|
return f;
|
2012-10-05 08:16:21 +08:00
|
|
|
}
|
|
|
|
}
|
2005-04-17 06:20:36 +08:00
|
|
|
return NULL;
|
|
|
|
}
|
|
|
|
|
|
|
|
/*
|
|
|
|
* Leave the top bit clear so we have tagspace for userland.
|
|
|
|
* The bottom 16 bits are the xmit tick for rexmit/rttavg processing.
|
|
|
|
* This driver reserves tag -1 to mean "unused frame."
|
|
|
|
*/
|
|
|
|
static int
|
2012-10-05 08:16:33 +08:00
|
|
|
newtag(struct aoedev *d)
|
2005-04-17 06:20:36 +08:00
|
|
|
{
|
|
|
|
register ulong n;
|
|
|
|
|
|
|
|
n = jiffies & 0xffff;
|
2012-10-05 08:16:33 +08:00
|
|
|
return n |= (++d->lasttag & 0x7fff) << 16;
|
2005-04-17 06:20:36 +08:00
|
|
|
}
|
|
|
|
|
2012-10-05 08:16:21 +08:00
|
|
|
static u32
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
aoehdr_atainit(struct aoedev *d, struct aoetgt *t, struct aoe_hdr *h)
|
2005-04-17 06:20:36 +08:00
|
|
|
{
|
2012-10-05 08:16:33 +08:00
|
|
|
u32 host_tag = newtag(d);
|
2005-04-17 06:20:36 +08:00
|
|
|
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
memcpy(h->src, t->ifp->nd->dev_addr, sizeof h->src);
|
|
|
|
memcpy(h->dst, t->addr, sizeof h->dst);
|
2005-04-19 13:00:20 +08:00
|
|
|
h->type = __constant_cpu_to_be16(ETH_P_AOE);
|
2005-04-17 06:20:36 +08:00
|
|
|
h->verfl = AOE_HVER;
|
2005-04-19 13:00:20 +08:00
|
|
|
h->major = cpu_to_be16(d->aoemajor);
|
2005-04-17 06:20:36 +08:00
|
|
|
h->minor = d->aoeminor;
|
|
|
|
h->cmd = AOECMD_ATA;
|
2005-04-19 13:00:20 +08:00
|
|
|
h->tag = cpu_to_be32(host_tag);
|
2005-04-17 06:20:36 +08:00
|
|
|
|
|
|
|
return host_tag;
|
|
|
|
}
|
|
|
|
|
2006-09-21 02:36:49 +08:00
|
|
|
static inline void
|
|
|
|
put_lba(struct aoe_atahdr *ah, sector_t lba)
|
|
|
|
{
|
|
|
|
ah->lba0 = lba;
|
|
|
|
ah->lba1 = lba >>= 8;
|
|
|
|
ah->lba2 = lba >>= 8;
|
|
|
|
ah->lba3 = lba >>= 8;
|
|
|
|
ah->lba4 = lba >>= 8;
|
|
|
|
ah->lba5 = lba >>= 8;
|
|
|
|
}
|
|
|
|
|
2012-10-05 08:16:27 +08:00
|
|
|
static struct aoeif *
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
ifrotate(struct aoetgt *t)
|
|
|
|
{
|
2012-10-05 08:16:27 +08:00
|
|
|
struct aoeif *ifp;
|
|
|
|
|
|
|
|
ifp = t->ifp;
|
|
|
|
ifp++;
|
|
|
|
if (ifp >= &t->ifs[NAOEIFS] || ifp->nd == NULL)
|
|
|
|
ifp = t->ifs;
|
|
|
|
if (ifp->nd == NULL)
|
|
|
|
return NULL;
|
|
|
|
return t->ifp = ifp;
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
}
|
|
|
|
|
aoe: dynamically allocate a capped number of skbs when necessary
What this Patch Does
Even before this recent series of 12 patches to 2.6.22-rc4, the aoe
driver was reusing a small set of skbs that were allocated once and
were only used for outbound AoE commands.
The network layer cannot be allowed to put_page on the data that is
still associated with a bio we haven't returned to the block layer,
so the aoe driver (even before the patch under discussion) is still
the owner of skbs that have been handed to the network layer for
transmission. We need to keep track of these skbs so that we can
free them, but by tracking them, we can also easily re-use them.
The new patch was a response to the behavior of certain network
drivers. We cannot reuse an skb that the network driver still has
in its transmit ring. Network drivers can defer transmit ring
cleanup and then use the state in the skb to determine how many data
segments to clean up in its transmit ring. The tg3 driver is one
driver that behaves in this way.
When the network driver defers cleanup of its transmit ring, the aoe
driver can find itself in a situation where it would like to send an
AoE command, and the AoE target is ready for more work, but the
network driver still has all of the pre-allocated skbs. In that
case, the new patch just calls alloc_skb, as you'd expect.
We don't want to get carried away, though. We try not to do
excessive allocation in the write path, so we cap the number of skbs
we dynamically allocate.
Probably calling it a "dynamic pool" is misleading. We were already
trying to use a small fixed-size set of pre-allocated skbs before
this patch, and this patch just provides a little headroom (with a
ceiling, though) to accomodate network drivers that hang onto skbs,
by allocating when needed. The d->skbpool_hd list of allocated skbs
is necessary so that we can free them later.
We didn't notice the need for this headroom until AoE targets got
fast enough.
Alternatives
If the network layer never did a put_page on the pages in the bio's
we get from the block layer, then it would be possible for us to
hand skbs to the network layer and forget about them, allowing the
network layer to free skbs itself (and thereby calling our own
skb->destructor callback function if we needed that). In that case
we could get rid of the pre-allocated skbs and also the
d->skbpool_hd, instead just calling alloc_skb every time we wanted
to transmit a packet. The slab allocator would effectively maintain
the list of skbs.
Besides a loss of CPU cache locality, the main concern with that
approach the danger that it would increase the likelihood of
deadlock when VM is trying to free pages by writing dirty data from
the page cache through the aoe driver out to persistent storage on
an AoE device. Right now we have a situation where we have
pre-allocation that corresponds to how much we use, which seems
ideal.
Of course, there's still the separate issue of receiving the packets
that tell us that a write has successfully completed on the AoE
target. When memory is low and VM is using AoE to flush dirty data
to free up pages, it would be perfect if there were a way for us to
register a fast callback that could recognize write command
completion responses. But I don't think the current problems with
the receive side of the situation are a justification for
exacerbating the problem on the transmit side.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:05 +08:00
|
|
|
static void
|
|
|
|
skb_pool_put(struct aoedev *d, struct sk_buff *skb)
|
|
|
|
{
|
2008-09-22 13:36:49 +08:00
|
|
|
__skb_queue_tail(&d->skbpool, skb);
|
aoe: dynamically allocate a capped number of skbs when necessary
What this Patch Does
Even before this recent series of 12 patches to 2.6.22-rc4, the aoe
driver was reusing a small set of skbs that were allocated once and
were only used for outbound AoE commands.
The network layer cannot be allowed to put_page on the data that is
still associated with a bio we haven't returned to the block layer,
so the aoe driver (even before the patch under discussion) is still
the owner of skbs that have been handed to the network layer for
transmission. We need to keep track of these skbs so that we can
free them, but by tracking them, we can also easily re-use them.
The new patch was a response to the behavior of certain network
drivers. We cannot reuse an skb that the network driver still has
in its transmit ring. Network drivers can defer transmit ring
cleanup and then use the state in the skb to determine how many data
segments to clean up in its transmit ring. The tg3 driver is one
driver that behaves in this way.
When the network driver defers cleanup of its transmit ring, the aoe
driver can find itself in a situation where it would like to send an
AoE command, and the AoE target is ready for more work, but the
network driver still has all of the pre-allocated skbs. In that
case, the new patch just calls alloc_skb, as you'd expect.
We don't want to get carried away, though. We try not to do
excessive allocation in the write path, so we cap the number of skbs
we dynamically allocate.
Probably calling it a "dynamic pool" is misleading. We were already
trying to use a small fixed-size set of pre-allocated skbs before
this patch, and this patch just provides a little headroom (with a
ceiling, though) to accomodate network drivers that hang onto skbs,
by allocating when needed. The d->skbpool_hd list of allocated skbs
is necessary so that we can free them later.
We didn't notice the need for this headroom until AoE targets got
fast enough.
Alternatives
If the network layer never did a put_page on the pages in the bio's
we get from the block layer, then it would be possible for us to
hand skbs to the network layer and forget about them, allowing the
network layer to free skbs itself (and thereby calling our own
skb->destructor callback function if we needed that). In that case
we could get rid of the pre-allocated skbs and also the
d->skbpool_hd, instead just calling alloc_skb every time we wanted
to transmit a packet. The slab allocator would effectively maintain
the list of skbs.
Besides a loss of CPU cache locality, the main concern with that
approach the danger that it would increase the likelihood of
deadlock when VM is trying to free pages by writing dirty data from
the page cache through the aoe driver out to persistent storage on
an AoE device. Right now we have a situation where we have
pre-allocation that corresponds to how much we use, which seems
ideal.
Of course, there's still the separate issue of receiving the packets
that tell us that a write has successfully completed on the AoE
target. When memory is low and VM is using AoE to flush dirty data
to free up pages, it would be perfect if there were a way for us to
register a fast callback that could recognize write command
completion responses. But I don't think the current problems with
the receive side of the situation are a justification for
exacerbating the problem on the transmit side.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:05 +08:00
|
|
|
}
|
|
|
|
|
|
|
|
static struct sk_buff *
|
|
|
|
skb_pool_get(struct aoedev *d)
|
|
|
|
{
|
2008-09-22 13:36:49 +08:00
|
|
|
struct sk_buff *skb = skb_peek(&d->skbpool);
|
aoe: dynamically allocate a capped number of skbs when necessary
What this Patch Does
Even before this recent series of 12 patches to 2.6.22-rc4, the aoe
driver was reusing a small set of skbs that were allocated once and
were only used for outbound AoE commands.
The network layer cannot be allowed to put_page on the data that is
still associated with a bio we haven't returned to the block layer,
so the aoe driver (even before the patch under discussion) is still
the owner of skbs that have been handed to the network layer for
transmission. We need to keep track of these skbs so that we can
free them, but by tracking them, we can also easily re-use them.
The new patch was a response to the behavior of certain network
drivers. We cannot reuse an skb that the network driver still has
in its transmit ring. Network drivers can defer transmit ring
cleanup and then use the state in the skb to determine how many data
segments to clean up in its transmit ring. The tg3 driver is one
driver that behaves in this way.
When the network driver defers cleanup of its transmit ring, the aoe
driver can find itself in a situation where it would like to send an
AoE command, and the AoE target is ready for more work, but the
network driver still has all of the pre-allocated skbs. In that
case, the new patch just calls alloc_skb, as you'd expect.
We don't want to get carried away, though. We try not to do
excessive allocation in the write path, so we cap the number of skbs
we dynamically allocate.
Probably calling it a "dynamic pool" is misleading. We were already
trying to use a small fixed-size set of pre-allocated skbs before
this patch, and this patch just provides a little headroom (with a
ceiling, though) to accomodate network drivers that hang onto skbs,
by allocating when needed. The d->skbpool_hd list of allocated skbs
is necessary so that we can free them later.
We didn't notice the need for this headroom until AoE targets got
fast enough.
Alternatives
If the network layer never did a put_page on the pages in the bio's
we get from the block layer, then it would be possible for us to
hand skbs to the network layer and forget about them, allowing the
network layer to free skbs itself (and thereby calling our own
skb->destructor callback function if we needed that). In that case
we could get rid of the pre-allocated skbs and also the
d->skbpool_hd, instead just calling alloc_skb every time we wanted
to transmit a packet. The slab allocator would effectively maintain
the list of skbs.
Besides a loss of CPU cache locality, the main concern with that
approach the danger that it would increase the likelihood of
deadlock when VM is trying to free pages by writing dirty data from
the page cache through the aoe driver out to persistent storage on
an AoE device. Right now we have a situation where we have
pre-allocation that corresponds to how much we use, which seems
ideal.
Of course, there's still the separate issue of receiving the packets
that tell us that a write has successfully completed on the AoE
target. When memory is low and VM is using AoE to flush dirty data
to free up pages, it would be perfect if there were a way for us to
register a fast callback that could recognize write command
completion responses. But I don't think the current problems with
the receive side of the situation are a justification for
exacerbating the problem on the transmit side.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:05 +08:00
|
|
|
|
|
|
|
if (skb && atomic_read(&skb_shinfo(skb)->dataref) == 1) {
|
2008-09-22 13:36:49 +08:00
|
|
|
__skb_unlink(skb, &d->skbpool);
|
aoe: dynamically allocate a capped number of skbs when necessary
What this Patch Does
Even before this recent series of 12 patches to 2.6.22-rc4, the aoe
driver was reusing a small set of skbs that were allocated once and
were only used for outbound AoE commands.
The network layer cannot be allowed to put_page on the data that is
still associated with a bio we haven't returned to the block layer,
so the aoe driver (even before the patch under discussion) is still
the owner of skbs that have been handed to the network layer for
transmission. We need to keep track of these skbs so that we can
free them, but by tracking them, we can also easily re-use them.
The new patch was a response to the behavior of certain network
drivers. We cannot reuse an skb that the network driver still has
in its transmit ring. Network drivers can defer transmit ring
cleanup and then use the state in the skb to determine how many data
segments to clean up in its transmit ring. The tg3 driver is one
driver that behaves in this way.
When the network driver defers cleanup of its transmit ring, the aoe
driver can find itself in a situation where it would like to send an
AoE command, and the AoE target is ready for more work, but the
network driver still has all of the pre-allocated skbs. In that
case, the new patch just calls alloc_skb, as you'd expect.
We don't want to get carried away, though. We try not to do
excessive allocation in the write path, so we cap the number of skbs
we dynamically allocate.
Probably calling it a "dynamic pool" is misleading. We were already
trying to use a small fixed-size set of pre-allocated skbs before
this patch, and this patch just provides a little headroom (with a
ceiling, though) to accomodate network drivers that hang onto skbs,
by allocating when needed. The d->skbpool_hd list of allocated skbs
is necessary so that we can free them later.
We didn't notice the need for this headroom until AoE targets got
fast enough.
Alternatives
If the network layer never did a put_page on the pages in the bio's
we get from the block layer, then it would be possible for us to
hand skbs to the network layer and forget about them, allowing the
network layer to free skbs itself (and thereby calling our own
skb->destructor callback function if we needed that). In that case
we could get rid of the pre-allocated skbs and also the
d->skbpool_hd, instead just calling alloc_skb every time we wanted
to transmit a packet. The slab allocator would effectively maintain
the list of skbs.
Besides a loss of CPU cache locality, the main concern with that
approach the danger that it would increase the likelihood of
deadlock when VM is trying to free pages by writing dirty data from
the page cache through the aoe driver out to persistent storage on
an AoE device. Right now we have a situation where we have
pre-allocation that corresponds to how much we use, which seems
ideal.
Of course, there's still the separate issue of receiving the packets
that tell us that a write has successfully completed on the AoE
target. When memory is low and VM is using AoE to flush dirty data
to free up pages, it would be perfect if there were a way for us to
register a fast callback that could recognize write command
completion responses. But I don't think the current problems with
the receive side of the situation are a justification for
exacerbating the problem on the transmit side.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:05 +08:00
|
|
|
return skb;
|
|
|
|
}
|
2008-09-22 13:36:49 +08:00
|
|
|
if (skb_queue_len(&d->skbpool) < NSKBPOOLMAX &&
|
|
|
|
(skb = new_skb(ETH_ZLEN)))
|
aoe: dynamically allocate a capped number of skbs when necessary
What this Patch Does
Even before this recent series of 12 patches to 2.6.22-rc4, the aoe
driver was reusing a small set of skbs that were allocated once and
were only used for outbound AoE commands.
The network layer cannot be allowed to put_page on the data that is
still associated with a bio we haven't returned to the block layer,
so the aoe driver (even before the patch under discussion) is still
the owner of skbs that have been handed to the network layer for
transmission. We need to keep track of these skbs so that we can
free them, but by tracking them, we can also easily re-use them.
The new patch was a response to the behavior of certain network
drivers. We cannot reuse an skb that the network driver still has
in its transmit ring. Network drivers can defer transmit ring
cleanup and then use the state in the skb to determine how many data
segments to clean up in its transmit ring. The tg3 driver is one
driver that behaves in this way.
When the network driver defers cleanup of its transmit ring, the aoe
driver can find itself in a situation where it would like to send an
AoE command, and the AoE target is ready for more work, but the
network driver still has all of the pre-allocated skbs. In that
case, the new patch just calls alloc_skb, as you'd expect.
We don't want to get carried away, though. We try not to do
excessive allocation in the write path, so we cap the number of skbs
we dynamically allocate.
Probably calling it a "dynamic pool" is misleading. We were already
trying to use a small fixed-size set of pre-allocated skbs before
this patch, and this patch just provides a little headroom (with a
ceiling, though) to accomodate network drivers that hang onto skbs,
by allocating when needed. The d->skbpool_hd list of allocated skbs
is necessary so that we can free them later.
We didn't notice the need for this headroom until AoE targets got
fast enough.
Alternatives
If the network layer never did a put_page on the pages in the bio's
we get from the block layer, then it would be possible for us to
hand skbs to the network layer and forget about them, allowing the
network layer to free skbs itself (and thereby calling our own
skb->destructor callback function if we needed that). In that case
we could get rid of the pre-allocated skbs and also the
d->skbpool_hd, instead just calling alloc_skb every time we wanted
to transmit a packet. The slab allocator would effectively maintain
the list of skbs.
Besides a loss of CPU cache locality, the main concern with that
approach the danger that it would increase the likelihood of
deadlock when VM is trying to free pages by writing dirty data from
the page cache through the aoe driver out to persistent storage on
an AoE device. Right now we have a situation where we have
pre-allocation that corresponds to how much we use, which seems
ideal.
Of course, there's still the separate issue of receiving the packets
that tell us that a write has successfully completed on the AoE
target. When memory is low and VM is using AoE to flush dirty data
to free up pages, it would be perfect if there were a way for us to
register a fast callback that could recognize write command
completion responses. But I don't think the current problems with
the receive side of the situation are a justification for
exacerbating the problem on the transmit side.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:05 +08:00
|
|
|
return skb;
|
2008-09-22 13:36:49 +08:00
|
|
|
|
aoe: dynamically allocate a capped number of skbs when necessary
What this Patch Does
Even before this recent series of 12 patches to 2.6.22-rc4, the aoe
driver was reusing a small set of skbs that were allocated once and
were only used for outbound AoE commands.
The network layer cannot be allowed to put_page on the data that is
still associated with a bio we haven't returned to the block layer,
so the aoe driver (even before the patch under discussion) is still
the owner of skbs that have been handed to the network layer for
transmission. We need to keep track of these skbs so that we can
free them, but by tracking them, we can also easily re-use them.
The new patch was a response to the behavior of certain network
drivers. We cannot reuse an skb that the network driver still has
in its transmit ring. Network drivers can defer transmit ring
cleanup and then use the state in the skb to determine how many data
segments to clean up in its transmit ring. The tg3 driver is one
driver that behaves in this way.
When the network driver defers cleanup of its transmit ring, the aoe
driver can find itself in a situation where it would like to send an
AoE command, and the AoE target is ready for more work, but the
network driver still has all of the pre-allocated skbs. In that
case, the new patch just calls alloc_skb, as you'd expect.
We don't want to get carried away, though. We try not to do
excessive allocation in the write path, so we cap the number of skbs
we dynamically allocate.
Probably calling it a "dynamic pool" is misleading. We were already
trying to use a small fixed-size set of pre-allocated skbs before
this patch, and this patch just provides a little headroom (with a
ceiling, though) to accomodate network drivers that hang onto skbs,
by allocating when needed. The d->skbpool_hd list of allocated skbs
is necessary so that we can free them later.
We didn't notice the need for this headroom until AoE targets got
fast enough.
Alternatives
If the network layer never did a put_page on the pages in the bio's
we get from the block layer, then it would be possible for us to
hand skbs to the network layer and forget about them, allowing the
network layer to free skbs itself (and thereby calling our own
skb->destructor callback function if we needed that). In that case
we could get rid of the pre-allocated skbs and also the
d->skbpool_hd, instead just calling alloc_skb every time we wanted
to transmit a packet. The slab allocator would effectively maintain
the list of skbs.
Besides a loss of CPU cache locality, the main concern with that
approach the danger that it would increase the likelihood of
deadlock when VM is trying to free pages by writing dirty data from
the page cache through the aoe driver out to persistent storage on
an AoE device. Right now we have a situation where we have
pre-allocation that corresponds to how much we use, which seems
ideal.
Of course, there's still the separate issue of receiving the packets
that tell us that a write has successfully completed on the AoE
target. When memory is low and VM is using AoE to flush dirty data
to free up pages, it would be perfect if there were a way for us to
register a fast callback that could recognize write command
completion responses. But I don't think the current problems with
the receive side of the situation are a justification for
exacerbating the problem on the transmit side.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:05 +08:00
|
|
|
return NULL;
|
|
|
|
}
|
|
|
|
|
2012-10-05 08:16:21 +08:00
|
|
|
void
|
|
|
|
aoe_freetframe(struct frame *f)
|
|
|
|
{
|
|
|
|
struct aoetgt *t;
|
|
|
|
|
|
|
|
t = f->t;
|
|
|
|
f->buf = NULL;
|
|
|
|
f->bv = NULL;
|
|
|
|
f->r_skb = NULL;
|
|
|
|
list_add(&f->head, &t->ffree);
|
|
|
|
}
|
|
|
|
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
static struct frame *
|
2012-10-05 08:16:21 +08:00
|
|
|
newtframe(struct aoedev *d, struct aoetgt *t)
|
2005-04-17 06:20:36 +08:00
|
|
|
{
|
2012-10-05 08:16:21 +08:00
|
|
|
struct frame *f;
|
aoe: dynamically allocate a capped number of skbs when necessary
What this Patch Does
Even before this recent series of 12 patches to 2.6.22-rc4, the aoe
driver was reusing a small set of skbs that were allocated once and
were only used for outbound AoE commands.
The network layer cannot be allowed to put_page on the data that is
still associated with a bio we haven't returned to the block layer,
so the aoe driver (even before the patch under discussion) is still
the owner of skbs that have been handed to the network layer for
transmission. We need to keep track of these skbs so that we can
free them, but by tracking them, we can also easily re-use them.
The new patch was a response to the behavior of certain network
drivers. We cannot reuse an skb that the network driver still has
in its transmit ring. Network drivers can defer transmit ring
cleanup and then use the state in the skb to determine how many data
segments to clean up in its transmit ring. The tg3 driver is one
driver that behaves in this way.
When the network driver defers cleanup of its transmit ring, the aoe
driver can find itself in a situation where it would like to send an
AoE command, and the AoE target is ready for more work, but the
network driver still has all of the pre-allocated skbs. In that
case, the new patch just calls alloc_skb, as you'd expect.
We don't want to get carried away, though. We try not to do
excessive allocation in the write path, so we cap the number of skbs
we dynamically allocate.
Probably calling it a "dynamic pool" is misleading. We were already
trying to use a small fixed-size set of pre-allocated skbs before
this patch, and this patch just provides a little headroom (with a
ceiling, though) to accomodate network drivers that hang onto skbs,
by allocating when needed. The d->skbpool_hd list of allocated skbs
is necessary so that we can free them later.
We didn't notice the need for this headroom until AoE targets got
fast enough.
Alternatives
If the network layer never did a put_page on the pages in the bio's
we get from the block layer, then it would be possible for us to
hand skbs to the network layer and forget about them, allowing the
network layer to free skbs itself (and thereby calling our own
skb->destructor callback function if we needed that). In that case
we could get rid of the pre-allocated skbs and also the
d->skbpool_hd, instead just calling alloc_skb every time we wanted
to transmit a packet. The slab allocator would effectively maintain
the list of skbs.
Besides a loss of CPU cache locality, the main concern with that
approach the danger that it would increase the likelihood of
deadlock when VM is trying to free pages by writing dirty data from
the page cache through the aoe driver out to persistent storage on
an AoE device. Right now we have a situation where we have
pre-allocation that corresponds to how much we use, which seems
ideal.
Of course, there's still the separate issue of receiving the packets
that tell us that a write has successfully completed on the AoE
target. When memory is low and VM is using AoE to flush dirty data
to free up pages, it would be perfect if there were a way for us to
register a fast callback that could recognize write command
completion responses. But I don't think the current problems with
the receive side of the situation are a justification for
exacerbating the problem on the transmit side.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:05 +08:00
|
|
|
struct sk_buff *skb;
|
2012-10-05 08:16:21 +08:00
|
|
|
struct list_head *pos;
|
|
|
|
|
|
|
|
if (list_empty(&t->ffree)) {
|
|
|
|
if (t->falloc >= NSKBPOOLMAX*2)
|
|
|
|
return NULL;
|
|
|
|
f = kcalloc(1, sizeof(*f), GFP_ATOMIC);
|
|
|
|
if (f == NULL)
|
|
|
|
return NULL;
|
|
|
|
t->falloc++;
|
|
|
|
f->t = t;
|
|
|
|
} else {
|
|
|
|
pos = t->ffree.next;
|
|
|
|
list_del(pos);
|
|
|
|
f = list_entry(pos, struct frame, head);
|
|
|
|
}
|
|
|
|
|
|
|
|
skb = f->skb;
|
|
|
|
if (skb == NULL) {
|
|
|
|
f->skb = skb = new_skb(ETH_ZLEN);
|
|
|
|
if (!skb) {
|
|
|
|
bail: aoe_freetframe(f);
|
|
|
|
return NULL;
|
|
|
|
}
|
|
|
|
}
|
|
|
|
|
|
|
|
if (atomic_read(&skb_shinfo(skb)->dataref) != 1) {
|
|
|
|
skb = skb_pool_get(d);
|
|
|
|
if (skb == NULL)
|
|
|
|
goto bail;
|
|
|
|
skb_pool_put(d, f->skb);
|
|
|
|
f->skb = skb;
|
|
|
|
}
|
|
|
|
|
|
|
|
skb->truesize -= skb->data_len;
|
|
|
|
skb_shinfo(skb)->nr_frags = skb->data_len = 0;
|
|
|
|
skb_trim(skb, 0);
|
|
|
|
return f;
|
|
|
|
}
|
|
|
|
|
|
|
|
static struct frame *
|
|
|
|
newframe(struct aoedev *d)
|
|
|
|
{
|
|
|
|
struct frame *f;
|
|
|
|
struct aoetgt *t, **tt;
|
|
|
|
int totout = 0;
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
|
|
|
|
if (d->targets[0] == NULL) { /* shouldn't happen, but I'm paranoid */
|
|
|
|
printk(KERN_ERR "aoe: NULL TARGETS!\n");
|
|
|
|
return NULL;
|
|
|
|
}
|
2012-10-05 08:16:21 +08:00
|
|
|
tt = d->tgt; /* last used target */
|
aoe: dynamically allocate a capped number of skbs when necessary
What this Patch Does
Even before this recent series of 12 patches to 2.6.22-rc4, the aoe
driver was reusing a small set of skbs that were allocated once and
were only used for outbound AoE commands.
The network layer cannot be allowed to put_page on the data that is
still associated with a bio we haven't returned to the block layer,
so the aoe driver (even before the patch under discussion) is still
the owner of skbs that have been handed to the network layer for
transmission. We need to keep track of these skbs so that we can
free them, but by tracking them, we can also easily re-use them.
The new patch was a response to the behavior of certain network
drivers. We cannot reuse an skb that the network driver still has
in its transmit ring. Network drivers can defer transmit ring
cleanup and then use the state in the skb to determine how many data
segments to clean up in its transmit ring. The tg3 driver is one
driver that behaves in this way.
When the network driver defers cleanup of its transmit ring, the aoe
driver can find itself in a situation where it would like to send an
AoE command, and the AoE target is ready for more work, but the
network driver still has all of the pre-allocated skbs. In that
case, the new patch just calls alloc_skb, as you'd expect.
We don't want to get carried away, though. We try not to do
excessive allocation in the write path, so we cap the number of skbs
we dynamically allocate.
Probably calling it a "dynamic pool" is misleading. We were already
trying to use a small fixed-size set of pre-allocated skbs before
this patch, and this patch just provides a little headroom (with a
ceiling, though) to accomodate network drivers that hang onto skbs,
by allocating when needed. The d->skbpool_hd list of allocated skbs
is necessary so that we can free them later.
We didn't notice the need for this headroom until AoE targets got
fast enough.
Alternatives
If the network layer never did a put_page on the pages in the bio's
we get from the block layer, then it would be possible for us to
hand skbs to the network layer and forget about them, allowing the
network layer to free skbs itself (and thereby calling our own
skb->destructor callback function if we needed that). In that case
we could get rid of the pre-allocated skbs and also the
d->skbpool_hd, instead just calling alloc_skb every time we wanted
to transmit a packet. The slab allocator would effectively maintain
the list of skbs.
Besides a loss of CPU cache locality, the main concern with that
approach the danger that it would increase the likelihood of
deadlock when VM is trying to free pages by writing dirty data from
the page cache through the aoe driver out to persistent storage on
an AoE device. Right now we have a situation where we have
pre-allocation that corresponds to how much we use, which seems
ideal.
Of course, there's still the separate issue of receiving the packets
that tell us that a write has successfully completed on the AoE
target. When memory is low and VM is using AoE to flush dirty data
to free up pages, it would be perfect if there were a way for us to
register a fast callback that could recognize write command
completion responses. But I don't think the current problems with
the receive side of the situation are a justification for
exacerbating the problem on the transmit side.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:05 +08:00
|
|
|
for (;;) {
|
2012-10-05 08:16:21 +08:00
|
|
|
tt++;
|
|
|
|
if (tt >= &d->targets[NTARGETS] || !*tt)
|
|
|
|
tt = d->targets;
|
|
|
|
t = *tt;
|
|
|
|
totout += t->nout;
|
|
|
|
if (t->nout < t->maxout
|
aoe: dynamically allocate a capped number of skbs when necessary
What this Patch Does
Even before this recent series of 12 patches to 2.6.22-rc4, the aoe
driver was reusing a small set of skbs that were allocated once and
were only used for outbound AoE commands.
The network layer cannot be allowed to put_page on the data that is
still associated with a bio we haven't returned to the block layer,
so the aoe driver (even before the patch under discussion) is still
the owner of skbs that have been handed to the network layer for
transmission. We need to keep track of these skbs so that we can
free them, but by tracking them, we can also easily re-use them.
The new patch was a response to the behavior of certain network
drivers. We cannot reuse an skb that the network driver still has
in its transmit ring. Network drivers can defer transmit ring
cleanup and then use the state in the skb to determine how many data
segments to clean up in its transmit ring. The tg3 driver is one
driver that behaves in this way.
When the network driver defers cleanup of its transmit ring, the aoe
driver can find itself in a situation where it would like to send an
AoE command, and the AoE target is ready for more work, but the
network driver still has all of the pre-allocated skbs. In that
case, the new patch just calls alloc_skb, as you'd expect.
We don't want to get carried away, though. We try not to do
excessive allocation in the write path, so we cap the number of skbs
we dynamically allocate.
Probably calling it a "dynamic pool" is misleading. We were already
trying to use a small fixed-size set of pre-allocated skbs before
this patch, and this patch just provides a little headroom (with a
ceiling, though) to accomodate network drivers that hang onto skbs,
by allocating when needed. The d->skbpool_hd list of allocated skbs
is necessary so that we can free them later.
We didn't notice the need for this headroom until AoE targets got
fast enough.
Alternatives
If the network layer never did a put_page on the pages in the bio's
we get from the block layer, then it would be possible for us to
hand skbs to the network layer and forget about them, allowing the
network layer to free skbs itself (and thereby calling our own
skb->destructor callback function if we needed that). In that case
we could get rid of the pre-allocated skbs and also the
d->skbpool_hd, instead just calling alloc_skb every time we wanted
to transmit a packet. The slab allocator would effectively maintain
the list of skbs.
Besides a loss of CPU cache locality, the main concern with that
approach the danger that it would increase the likelihood of
deadlock when VM is trying to free pages by writing dirty data from
the page cache through the aoe driver out to persistent storage on
an AoE device. Right now we have a situation where we have
pre-allocation that corresponds to how much we use, which seems
ideal.
Of course, there's still the separate issue of receiving the packets
that tell us that a write has successfully completed on the AoE
target. When memory is low and VM is using AoE to flush dirty data
to free up pages, it would be perfect if there were a way for us to
register a fast callback that could recognize write command
completion responses. But I don't think the current problems with
the receive side of the situation are a justification for
exacerbating the problem on the transmit side.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:05 +08:00
|
|
|
&& t != d->htgt
|
2012-10-05 08:16:21 +08:00
|
|
|
&& t->ifp->nd) {
|
|
|
|
f = newtframe(d, t);
|
|
|
|
if (f) {
|
|
|
|
ifrotate(t);
|
2012-10-05 08:16:27 +08:00
|
|
|
d->tgt = tt;
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
return f;
|
|
|
|
}
|
|
|
|
}
|
2012-10-05 08:16:21 +08:00
|
|
|
if (tt == d->tgt) /* we've looped and found nada */
|
aoe: dynamically allocate a capped number of skbs when necessary
What this Patch Does
Even before this recent series of 12 patches to 2.6.22-rc4, the aoe
driver was reusing a small set of skbs that were allocated once and
were only used for outbound AoE commands.
The network layer cannot be allowed to put_page on the data that is
still associated with a bio we haven't returned to the block layer,
so the aoe driver (even before the patch under discussion) is still
the owner of skbs that have been handed to the network layer for
transmission. We need to keep track of these skbs so that we can
free them, but by tracking them, we can also easily re-use them.
The new patch was a response to the behavior of certain network
drivers. We cannot reuse an skb that the network driver still has
in its transmit ring. Network drivers can defer transmit ring
cleanup and then use the state in the skb to determine how many data
segments to clean up in its transmit ring. The tg3 driver is one
driver that behaves in this way.
When the network driver defers cleanup of its transmit ring, the aoe
driver can find itself in a situation where it would like to send an
AoE command, and the AoE target is ready for more work, but the
network driver still has all of the pre-allocated skbs. In that
case, the new patch just calls alloc_skb, as you'd expect.
We don't want to get carried away, though. We try not to do
excessive allocation in the write path, so we cap the number of skbs
we dynamically allocate.
Probably calling it a "dynamic pool" is misleading. We were already
trying to use a small fixed-size set of pre-allocated skbs before
this patch, and this patch just provides a little headroom (with a
ceiling, though) to accomodate network drivers that hang onto skbs,
by allocating when needed. The d->skbpool_hd list of allocated skbs
is necessary so that we can free them later.
We didn't notice the need for this headroom until AoE targets got
fast enough.
Alternatives
If the network layer never did a put_page on the pages in the bio's
we get from the block layer, then it would be possible for us to
hand skbs to the network layer and forget about them, allowing the
network layer to free skbs itself (and thereby calling our own
skb->destructor callback function if we needed that). In that case
we could get rid of the pre-allocated skbs and also the
d->skbpool_hd, instead just calling alloc_skb every time we wanted
to transmit a packet. The slab allocator would effectively maintain
the list of skbs.
Besides a loss of CPU cache locality, the main concern with that
approach the danger that it would increase the likelihood of
deadlock when VM is trying to free pages by writing dirty data from
the page cache through the aoe driver out to persistent storage on
an AoE device. Right now we have a situation where we have
pre-allocation that corresponds to how much we use, which seems
ideal.
Of course, there's still the separate issue of receiving the packets
that tell us that a write has successfully completed on the AoE
target. When memory is low and VM is using AoE to flush dirty data
to free up pages, it would be perfect if there were a way for us to
register a fast callback that could recognize write command
completion responses. But I don't think the current problems with
the receive side of the situation are a justification for
exacerbating the problem on the transmit side.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:05 +08:00
|
|
|
break;
|
2012-10-05 08:16:21 +08:00
|
|
|
}
|
|
|
|
if (totout == 0) {
|
|
|
|
d->kicked++;
|
|
|
|
d->flags |= DEVFL_KICKME;
|
aoe: dynamically allocate a capped number of skbs when necessary
What this Patch Does
Even before this recent series of 12 patches to 2.6.22-rc4, the aoe
driver was reusing a small set of skbs that were allocated once and
were only used for outbound AoE commands.
The network layer cannot be allowed to put_page on the data that is
still associated with a bio we haven't returned to the block layer,
so the aoe driver (even before the patch under discussion) is still
the owner of skbs that have been handed to the network layer for
transmission. We need to keep track of these skbs so that we can
free them, but by tracking them, we can also easily re-use them.
The new patch was a response to the behavior of certain network
drivers. We cannot reuse an skb that the network driver still has
in its transmit ring. Network drivers can defer transmit ring
cleanup and then use the state in the skb to determine how many data
segments to clean up in its transmit ring. The tg3 driver is one
driver that behaves in this way.
When the network driver defers cleanup of its transmit ring, the aoe
driver can find itself in a situation where it would like to send an
AoE command, and the AoE target is ready for more work, but the
network driver still has all of the pre-allocated skbs. In that
case, the new patch just calls alloc_skb, as you'd expect.
We don't want to get carried away, though. We try not to do
excessive allocation in the write path, so we cap the number of skbs
we dynamically allocate.
Probably calling it a "dynamic pool" is misleading. We were already
trying to use a small fixed-size set of pre-allocated skbs before
this patch, and this patch just provides a little headroom (with a
ceiling, though) to accomodate network drivers that hang onto skbs,
by allocating when needed. The d->skbpool_hd list of allocated skbs
is necessary so that we can free them later.
We didn't notice the need for this headroom until AoE targets got
fast enough.
Alternatives
If the network layer never did a put_page on the pages in the bio's
we get from the block layer, then it would be possible for us to
hand skbs to the network layer and forget about them, allowing the
network layer to free skbs itself (and thereby calling our own
skb->destructor callback function if we needed that). In that case
we could get rid of the pre-allocated skbs and also the
d->skbpool_hd, instead just calling alloc_skb every time we wanted
to transmit a packet. The slab allocator would effectively maintain
the list of skbs.
Besides a loss of CPU cache locality, the main concern with that
approach the danger that it would increase the likelihood of
deadlock when VM is trying to free pages by writing dirty data from
the page cache through the aoe driver out to persistent storage on
an AoE device. Right now we have a situation where we have
pre-allocation that corresponds to how much we use, which seems
ideal.
Of course, there's still the separate issue of receiving the packets
that tell us that a write has successfully completed on the AoE
target. When memory is low and VM is using AoE to flush dirty data
to free up pages, it would be perfect if there were a way for us to
register a fast callback that could recognize write command
completion responses. But I don't think the current problems with
the receive side of the situation are a justification for
exacerbating the problem on the transmit side.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:05 +08:00
|
|
|
}
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
return NULL;
|
|
|
|
}
|
|
|
|
|
2012-10-05 08:16:20 +08:00
|
|
|
static void
|
|
|
|
skb_fillup(struct sk_buff *skb, struct bio_vec *bv, ulong off, ulong cnt)
|
|
|
|
{
|
|
|
|
int frag = 0;
|
|
|
|
ulong fcnt;
|
|
|
|
loop:
|
|
|
|
fcnt = bv->bv_len - (off - bv->bv_offset);
|
|
|
|
if (fcnt > cnt)
|
|
|
|
fcnt = cnt;
|
|
|
|
skb_fill_page_desc(skb, frag++, bv->bv_page, off, fcnt);
|
|
|
|
cnt -= fcnt;
|
|
|
|
if (cnt <= 0)
|
|
|
|
return;
|
|
|
|
bv++;
|
|
|
|
off = bv->bv_offset;
|
|
|
|
goto loop;
|
|
|
|
}
|
|
|
|
|
2012-10-05 08:16:21 +08:00
|
|
|
static void
|
|
|
|
fhash(struct frame *f)
|
|
|
|
{
|
2012-10-05 08:16:33 +08:00
|
|
|
struct aoedev *d = f->t->d;
|
2012-10-05 08:16:21 +08:00
|
|
|
u32 n;
|
|
|
|
|
|
|
|
n = f->tag % NFACTIVE;
|
2012-10-05 08:16:33 +08:00
|
|
|
list_add_tail(&f->head, &d->factive[n]);
|
2012-10-05 08:16:21 +08:00
|
|
|
}
|
|
|
|
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
static int
|
|
|
|
aoecmd_ata_rw(struct aoedev *d)
|
|
|
|
{
|
|
|
|
struct frame *f;
|
2005-04-17 06:20:36 +08:00
|
|
|
struct aoe_hdr *h;
|
|
|
|
struct aoe_atahdr *ah;
|
|
|
|
struct buf *buf;
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
struct aoetgt *t;
|
2005-04-17 06:20:36 +08:00
|
|
|
struct sk_buff *skb;
|
2012-10-05 08:16:23 +08:00
|
|
|
struct sk_buff_head queue;
|
2012-10-05 08:16:20 +08:00
|
|
|
ulong bcnt, fbcnt;
|
2005-04-17 06:20:36 +08:00
|
|
|
char writebit, extbit;
|
|
|
|
|
|
|
|
writebit = 0x10;
|
|
|
|
extbit = 0x4;
|
|
|
|
|
2012-10-05 08:16:23 +08:00
|
|
|
buf = nextbuf(d);
|
|
|
|
if (buf == NULL)
|
|
|
|
return 0;
|
2012-10-05 08:16:21 +08:00
|
|
|
f = newframe(d);
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
if (f == NULL)
|
|
|
|
return 0;
|
|
|
|
t = *d->tgt;
|
2012-10-05 08:16:27 +08:00
|
|
|
bcnt = d->maxbcnt;
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
if (bcnt == 0)
|
|
|
|
bcnt = DEFAULTBCNT;
|
2012-10-05 08:16:20 +08:00
|
|
|
if (bcnt > buf->resid)
|
|
|
|
bcnt = buf->resid;
|
|
|
|
fbcnt = bcnt;
|
|
|
|
f->bv = buf->bv;
|
|
|
|
f->bv_off = f->bv->bv_offset + (f->bv->bv_len - buf->bv_resid);
|
|
|
|
do {
|
|
|
|
if (fbcnt < buf->bv_resid) {
|
|
|
|
buf->bv_resid -= fbcnt;
|
|
|
|
buf->resid -= fbcnt;
|
|
|
|
break;
|
|
|
|
}
|
|
|
|
fbcnt -= buf->bv_resid;
|
|
|
|
buf->resid -= buf->bv_resid;
|
|
|
|
if (buf->resid == 0) {
|
2012-10-05 08:16:23 +08:00
|
|
|
d->ip.buf = NULL;
|
2012-10-05 08:16:20 +08:00
|
|
|
break;
|
|
|
|
}
|
|
|
|
buf->bv++;
|
|
|
|
buf->bv_resid = buf->bv->bv_len;
|
|
|
|
WARN_ON(buf->bv_resid == 0);
|
|
|
|
} while (fbcnt);
|
|
|
|
|
2005-04-17 06:20:36 +08:00
|
|
|
/* initialize the headers & frame */
|
2006-09-21 02:36:49 +08:00
|
|
|
skb = f->skb;
|
2007-10-17 14:27:03 +08:00
|
|
|
h = (struct aoe_hdr *) skb_mac_header(skb);
|
2005-04-17 06:20:36 +08:00
|
|
|
ah = (struct aoe_atahdr *) (h+1);
|
2006-12-22 17:09:21 +08:00
|
|
|
skb_put(skb, sizeof *h + sizeof *ah);
|
|
|
|
memset(h, 0, skb->len);
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
f->tag = aoehdr_atainit(d, t, h);
|
2012-10-05 08:16:21 +08:00
|
|
|
fhash(f);
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
t->nout++;
|
2005-04-17 06:20:36 +08:00
|
|
|
f->waited = 0;
|
|
|
|
f->buf = buf;
|
2006-09-21 02:36:49 +08:00
|
|
|
f->bcnt = bcnt;
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
f->lba = buf->sector;
|
2005-04-17 06:20:36 +08:00
|
|
|
|
|
|
|
/* set up ata header */
|
|
|
|
ah->scnt = bcnt >> 9;
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
put_lba(ah, buf->sector);
|
2005-04-17 06:20:36 +08:00
|
|
|
if (d->flags & DEVFL_EXT) {
|
|
|
|
ah->aflags |= AOEAFL_EXT;
|
|
|
|
} else {
|
|
|
|
extbit = 0;
|
|
|
|
ah->lba3 &= 0x0f;
|
|
|
|
ah->lba3 |= 0xe0; /* LBA bit + obsolete 0xa0 */
|
|
|
|
}
|
|
|
|
if (bio_data_dir(buf->bio) == WRITE) {
|
2012-10-05 08:16:20 +08:00
|
|
|
skb_fillup(skb, f->bv, f->bv_off, bcnt);
|
2005-04-17 06:20:36 +08:00
|
|
|
ah->aflags |= AOEAFL_WRITE;
|
2006-09-21 02:36:49 +08:00
|
|
|
skb->len += bcnt;
|
|
|
|
skb->data_len = bcnt;
|
2012-10-05 08:16:20 +08:00
|
|
|
skb->truesize += bcnt;
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
t->wpkts++;
|
2005-04-17 06:20:36 +08:00
|
|
|
} else {
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
t->rpkts++;
|
2005-04-17 06:20:36 +08:00
|
|
|
writebit = 0;
|
|
|
|
}
|
|
|
|
|
2009-04-02 03:42:24 +08:00
|
|
|
ah->cmdstat = ATA_CMD_PIO_READ | writebit | extbit;
|
2005-04-17 06:20:36 +08:00
|
|
|
|
|
|
|
/* mark all tracking fields and load out */
|
|
|
|
buf->nframesout += 1;
|
|
|
|
buf->sector += bcnt >> 9;
|
|
|
|
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
skb->dev = t->ifp->nd;
|
2006-09-21 02:36:49 +08:00
|
|
|
skb = skb_clone(skb, GFP_ATOMIC);
|
2012-10-05 08:16:23 +08:00
|
|
|
if (skb) {
|
|
|
|
__skb_queue_head_init(&queue);
|
|
|
|
__skb_queue_tail(&queue, skb);
|
|
|
|
aoenet_xmit(&queue);
|
|
|
|
}
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
return 1;
|
2005-04-17 06:20:36 +08:00
|
|
|
}
|
|
|
|
|
2006-01-20 02:46:19 +08:00
|
|
|
/* some callers cannot sleep, and they can call this function,
|
|
|
|
* transmitting the packets later, when interrupts are on
|
|
|
|
*/
|
2008-09-22 13:36:49 +08:00
|
|
|
static void
|
|
|
|
aoecmd_cfg_pkts(ushort aoemajor, unsigned char aoeminor, struct sk_buff_head *queue)
|
2006-01-20 02:46:19 +08:00
|
|
|
{
|
|
|
|
struct aoe_hdr *h;
|
|
|
|
struct aoe_cfghdr *ch;
|
2008-09-22 13:36:49 +08:00
|
|
|
struct sk_buff *skb;
|
2006-01-20 02:46:19 +08:00
|
|
|
struct net_device *ifp;
|
|
|
|
|
2010-10-29 09:15:29 +08:00
|
|
|
rcu_read_lock();
|
|
|
|
for_each_netdev_rcu(&init_net, ifp) {
|
2006-01-20 02:46:19 +08:00
|
|
|
dev_hold(ifp);
|
|
|
|
if (!is_aoe_netif(ifp))
|
2007-05-04 06:13:45 +08:00
|
|
|
goto cont;
|
2006-01-20 02:46:19 +08:00
|
|
|
|
2006-09-21 02:36:49 +08:00
|
|
|
skb = new_skb(sizeof *h + sizeof *ch);
|
2006-01-20 02:46:19 +08:00
|
|
|
if (skb == NULL) {
|
2006-09-21 02:36:51 +08:00
|
|
|
printk(KERN_INFO "aoe: skb alloc failure\n");
|
2007-05-04 06:13:45 +08:00
|
|
|
goto cont;
|
2006-01-20 02:46:19 +08:00
|
|
|
}
|
2006-12-22 17:09:21 +08:00
|
|
|
skb_put(skb, sizeof *h + sizeof *ch);
|
2006-09-21 02:36:49 +08:00
|
|
|
skb->dev = ifp;
|
2008-09-22 13:36:49 +08:00
|
|
|
__skb_queue_tail(queue, skb);
|
2007-10-17 14:27:03 +08:00
|
|
|
h = (struct aoe_hdr *) skb_mac_header(skb);
|
2006-01-20 02:46:19 +08:00
|
|
|
memset(h, 0, sizeof *h + sizeof *ch);
|
|
|
|
|
|
|
|
memset(h->dst, 0xff, sizeof h->dst);
|
|
|
|
memcpy(h->src, ifp->dev_addr, sizeof h->src);
|
|
|
|
h->type = __constant_cpu_to_be16(ETH_P_AOE);
|
|
|
|
h->verfl = AOE_HVER;
|
|
|
|
h->major = cpu_to_be16(aoemajor);
|
|
|
|
h->minor = aoeminor;
|
|
|
|
h->cmd = AOECMD_CFG;
|
|
|
|
|
2007-05-04 06:13:45 +08:00
|
|
|
cont:
|
|
|
|
dev_put(ifp);
|
2006-01-20 02:46:19 +08:00
|
|
|
}
|
2010-10-29 09:15:29 +08:00
|
|
|
rcu_read_unlock();
|
2006-01-20 02:46:19 +08:00
|
|
|
}
|
|
|
|
|
2005-04-17 06:20:36 +08:00
|
|
|
static void
|
2012-10-05 08:16:21 +08:00
|
|
|
resend(struct aoedev *d, struct frame *f)
|
2005-04-17 06:20:36 +08:00
|
|
|
{
|
|
|
|
struct sk_buff *skb;
|
2012-10-05 08:16:23 +08:00
|
|
|
struct sk_buff_head queue;
|
2005-04-17 06:20:36 +08:00
|
|
|
struct aoe_hdr *h;
|
2006-09-21 02:36:49 +08:00
|
|
|
struct aoe_atahdr *ah;
|
2012-10-05 08:16:21 +08:00
|
|
|
struct aoetgt *t;
|
2005-04-17 06:20:36 +08:00
|
|
|
char buf[128];
|
|
|
|
u32 n;
|
|
|
|
|
2012-10-05 08:16:21 +08:00
|
|
|
t = f->t;
|
2012-10-05 08:16:33 +08:00
|
|
|
n = newtag(d);
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
skb = f->skb;
|
2012-10-05 08:16:27 +08:00
|
|
|
if (ifrotate(t) == NULL) {
|
|
|
|
/* probably can't happen, but set it up to fail anyway */
|
|
|
|
pr_info("aoe: resend: no interfaces to rotate to.\n");
|
|
|
|
ktcomplete(f, NULL);
|
|
|
|
return;
|
|
|
|
}
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
h = (struct aoe_hdr *) skb_mac_header(skb);
|
|
|
|
ah = (struct aoe_atahdr *) (h+1);
|
2005-04-17 06:20:36 +08:00
|
|
|
|
|
|
|
snprintf(buf, sizeof buf,
|
2008-11-25 16:40:37 +08:00
|
|
|
"%15s e%ld.%d oldtag=%08x@%08lx newtag=%08x s=%pm d=%pm nout=%d\n",
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
"retransmit", d->aoemajor, d->aoeminor, f->tag, jiffies, n,
|
2008-11-25 16:40:37 +08:00
|
|
|
h->src, h->dst, t->nout);
|
2005-04-17 06:20:36 +08:00
|
|
|
aoechr_error(buf);
|
|
|
|
|
|
|
|
f->tag = n;
|
2012-10-05 08:16:21 +08:00
|
|
|
fhash(f);
|
2005-04-19 13:00:20 +08:00
|
|
|
h->tag = cpu_to_be32(n);
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
memcpy(h->dst, t->addr, sizeof h->dst);
|
|
|
|
memcpy(h->src, t->ifp->nd->dev_addr, sizeof h->src);
|
|
|
|
|
|
|
|
skb->dev = t->ifp->nd;
|
2006-09-21 02:36:49 +08:00
|
|
|
skb = skb_clone(skb, GFP_ATOMIC);
|
|
|
|
if (skb == NULL)
|
|
|
|
return;
|
2012-10-05 08:16:23 +08:00
|
|
|
__skb_queue_head_init(&queue);
|
|
|
|
__skb_queue_tail(&queue, skb);
|
|
|
|
aoenet_xmit(&queue);
|
2005-04-17 06:20:36 +08:00
|
|
|
}
|
|
|
|
|
|
|
|
static int
|
2012-10-05 08:16:21 +08:00
|
|
|
tsince(u32 tag)
|
2005-04-17 06:20:36 +08:00
|
|
|
{
|
|
|
|
int n;
|
|
|
|
|
|
|
|
n = jiffies & 0xffff;
|
|
|
|
n -= tag & 0xffff;
|
|
|
|
if (n < 0)
|
|
|
|
n += 1<<16;
|
|
|
|
return n;
|
|
|
|
}
|
|
|
|
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
static struct aoeif *
|
|
|
|
getif(struct aoetgt *t, struct net_device *nd)
|
|
|
|
{
|
|
|
|
struct aoeif *p, *e;
|
|
|
|
|
|
|
|
p = t->ifs;
|
|
|
|
e = p + NAOEIFS;
|
|
|
|
for (; p < e; p++)
|
|
|
|
if (p->nd == nd)
|
|
|
|
return p;
|
|
|
|
return NULL;
|
|
|
|
}
|
|
|
|
|
|
|
|
static void
|
|
|
|
ejectif(struct aoetgt *t, struct aoeif *ifp)
|
|
|
|
{
|
|
|
|
struct aoeif *e;
|
2012-10-05 08:16:34 +08:00
|
|
|
struct net_device *nd;
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
ulong n;
|
|
|
|
|
2012-10-05 08:16:34 +08:00
|
|
|
nd = ifp->nd;
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
e = t->ifs + NAOEIFS - 1;
|
|
|
|
n = (e - ifp) * sizeof *ifp;
|
|
|
|
memmove(ifp, ifp+1, n);
|
|
|
|
e->nd = NULL;
|
2012-10-05 08:16:34 +08:00
|
|
|
dev_put(nd);
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
}
|
|
|
|
|
|
|
|
static int
|
|
|
|
sthtith(struct aoedev *d)
|
|
|
|
{
|
2012-10-05 08:16:21 +08:00
|
|
|
struct frame *f, *nf;
|
|
|
|
struct list_head *nx, *pos, *head;
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
struct sk_buff *skb;
|
2012-10-05 08:16:21 +08:00
|
|
|
struct aoetgt *ht = d->htgt;
|
|
|
|
int i;
|
|
|
|
|
|
|
|
for (i = 0; i < NFACTIVE; i++) {
|
2012-10-05 08:16:33 +08:00
|
|
|
head = &d->factive[i];
|
2012-10-05 08:16:21 +08:00
|
|
|
list_for_each_safe(pos, nx, head) {
|
|
|
|
f = list_entry(pos, struct frame, head);
|
2012-10-05 08:16:33 +08:00
|
|
|
if (f->t != ht)
|
|
|
|
continue;
|
|
|
|
|
2012-10-05 08:16:21 +08:00
|
|
|
nf = newframe(d);
|
|
|
|
if (!nf)
|
|
|
|
return 0;
|
|
|
|
|
|
|
|
/* remove frame from active list */
|
|
|
|
list_del(pos);
|
|
|
|
|
|
|
|
/* reassign all pertinent bits to new outbound frame */
|
|
|
|
skb = nf->skb;
|
|
|
|
nf->skb = f->skb;
|
|
|
|
nf->buf = f->buf;
|
|
|
|
nf->bcnt = f->bcnt;
|
|
|
|
nf->lba = f->lba;
|
|
|
|
nf->bv = f->bv;
|
|
|
|
nf->bv_off = f->bv_off;
|
|
|
|
nf->waited = 0;
|
|
|
|
f->skb = skb;
|
|
|
|
aoe_freetframe(f);
|
|
|
|
ht->nout--;
|
|
|
|
nf->t->nout++;
|
|
|
|
resend(d, nf);
|
|
|
|
}
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
}
|
2012-10-05 08:16:27 +08:00
|
|
|
/* We've cleaned up the outstanding so take away his
|
|
|
|
* interfaces so he won't be used. We should remove him from
|
|
|
|
* the target array here, but cleaning up a target is
|
|
|
|
* involved. PUNT!
|
|
|
|
*/
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
memset(ht->ifs, 0, sizeof ht->ifs);
|
|
|
|
d->htgt = NULL;
|
|
|
|
return 1;
|
|
|
|
}
|
|
|
|
|
|
|
|
static inline unsigned char
|
|
|
|
ata_scnt(unsigned char *packet) {
|
|
|
|
struct aoe_hdr *h;
|
|
|
|
struct aoe_atahdr *ah;
|
|
|
|
|
|
|
|
h = (struct aoe_hdr *) packet;
|
|
|
|
ah = (struct aoe_atahdr *) (h+1);
|
|
|
|
return ah->scnt;
|
|
|
|
}
|
|
|
|
|
2005-04-17 06:20:36 +08:00
|
|
|
static void
|
|
|
|
rexmit_timer(ulong vp)
|
|
|
|
{
|
|
|
|
struct aoedev *d;
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
struct aoetgt *t, **tt, **te;
|
|
|
|
struct aoeif *ifp;
|
2012-10-05 08:16:21 +08:00
|
|
|
struct frame *f;
|
|
|
|
struct list_head *head, *pos, *nx;
|
|
|
|
LIST_HEAD(flist);
|
2005-04-17 06:20:36 +08:00
|
|
|
register long timeout;
|
|
|
|
ulong flags, n;
|
2012-10-05 08:16:21 +08:00
|
|
|
int i;
|
2005-04-17 06:20:36 +08:00
|
|
|
|
|
|
|
d = (struct aoedev *) vp;
|
|
|
|
|
|
|
|
/* timeout is always ~150% of the moving average */
|
|
|
|
timeout = d->rttavg;
|
|
|
|
timeout += timeout >> 1;
|
|
|
|
|
|
|
|
spin_lock_irqsave(&d->lock, flags);
|
|
|
|
|
|
|
|
if (d->flags & DEVFL_TKILL) {
|
2006-01-26 02:54:44 +08:00
|
|
|
spin_unlock_irqrestore(&d->lock, flags);
|
2005-04-17 06:20:36 +08:00
|
|
|
return;
|
|
|
|
}
|
2012-10-05 08:16:21 +08:00
|
|
|
|
|
|
|
/* collect all frames to rexmit into flist */
|
2012-10-05 08:16:33 +08:00
|
|
|
for (i = 0; i < NFACTIVE; i++) {
|
|
|
|
head = &d->factive[i];
|
|
|
|
list_for_each_safe(pos, nx, head) {
|
|
|
|
f = list_entry(pos, struct frame, head);
|
|
|
|
if (tsince(f->tag) < timeout)
|
|
|
|
break; /* end of expired frames */
|
|
|
|
/* move to flist for later processing */
|
|
|
|
list_move_tail(pos, &flist);
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
}
|
2012-10-05 08:16:33 +08:00
|
|
|
}
|
|
|
|
/* window check */
|
|
|
|
tt = d->targets;
|
|
|
|
te = tt + d->ntargets;
|
|
|
|
for (; tt < te && (t = *tt); tt++) {
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
if (t->nout == t->maxout
|
|
|
|
&& t->maxout < t->nframes
|
|
|
|
&& (jiffies - t->lastwadj)/HZ > 10) {
|
|
|
|
t->maxout++;
|
|
|
|
t->lastwadj = jiffies;
|
2005-04-17 06:20:36 +08:00
|
|
|
}
|
|
|
|
}
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
|
2012-10-05 08:16:23 +08:00
|
|
|
if (!list_empty(&flist)) { /* retransmissions necessary */
|
|
|
|
n = d->rttavg <<= 1;
|
|
|
|
if (n > MAXTIMER)
|
|
|
|
d->rttavg = MAXTIMER;
|
|
|
|
}
|
|
|
|
|
2012-10-05 08:16:21 +08:00
|
|
|
/* process expired frames */
|
|
|
|
while (!list_empty(&flist)) {
|
|
|
|
pos = flist.next;
|
|
|
|
f = list_entry(pos, struct frame, head);
|
|
|
|
n = f->waited += timeout;
|
|
|
|
n /= HZ;
|
|
|
|
if (n > aoe_deadsecs) {
|
|
|
|
/* Waited too long. Device failure.
|
|
|
|
* Hang all frames on first hash bucket for downdev
|
|
|
|
* to clean up.
|
|
|
|
*/
|
2012-10-05 08:16:33 +08:00
|
|
|
list_splice(&flist, &d->factive[0]);
|
2012-10-05 08:16:21 +08:00
|
|
|
aoedev_downdev(d);
|
|
|
|
break;
|
|
|
|
}
|
|
|
|
list_del(pos);
|
|
|
|
|
|
|
|
t = f->t;
|
2012-10-05 08:16:29 +08:00
|
|
|
if (n > aoe_deadsecs/2)
|
|
|
|
d->htgt = t; /* see if another target can help */
|
|
|
|
|
2012-10-05 08:16:21 +08:00
|
|
|
if (t->nout == t->maxout) {
|
|
|
|
if (t->maxout > 1)
|
|
|
|
t->maxout--;
|
|
|
|
t->lastwadj = jiffies;
|
|
|
|
}
|
|
|
|
|
|
|
|
ifp = getif(t, f->skb->dev);
|
|
|
|
if (ifp && ++ifp->lost > (t->nframes << 1)
|
|
|
|
&& (ifp != t->ifs || t->ifs[1].nd)) {
|
|
|
|
ejectif(t, ifp);
|
|
|
|
ifp = NULL;
|
|
|
|
}
|
|
|
|
resend(d, f);
|
|
|
|
}
|
|
|
|
|
2012-10-05 08:16:23 +08:00
|
|
|
if ((d->flags & DEVFL_KICKME || d->htgt) && d->blkq) {
|
2006-09-21 02:36:49 +08:00
|
|
|
d->flags &= ~DEVFL_KICKME;
|
2012-10-05 08:16:23 +08:00
|
|
|
d->blkq->request_fn(d->blkq);
|
2006-09-21 02:36:49 +08:00
|
|
|
}
|
2005-04-17 06:20:36 +08:00
|
|
|
|
|
|
|
d->timer.expires = jiffies + TIMERTICK;
|
|
|
|
add_timer(&d->timer);
|
|
|
|
|
|
|
|
spin_unlock_irqrestore(&d->lock, flags);
|
2012-10-05 08:16:23 +08:00
|
|
|
}
|
2005-04-17 06:20:36 +08:00
|
|
|
|
2012-10-05 08:16:23 +08:00
|
|
|
static unsigned long
|
|
|
|
rqbiocnt(struct request *r)
|
|
|
|
{
|
|
|
|
struct bio *bio;
|
|
|
|
unsigned long n = 0;
|
|
|
|
|
|
|
|
__rq_for_each_bio(bio, r)
|
|
|
|
n++;
|
|
|
|
return n;
|
|
|
|
}
|
|
|
|
|
|
|
|
/* This can be removed if we are certain that no users of the block
|
|
|
|
* layer will ever use zero-count pages in bios. Otherwise we have to
|
|
|
|
* protect against the put_page sometimes done by the network layer.
|
|
|
|
*
|
|
|
|
* See http://oss.sgi.com/archives/xfs/2007-01/msg00594.html for
|
|
|
|
* discussion.
|
|
|
|
*
|
|
|
|
* We cannot use get_page in the workaround, because it insists on a
|
|
|
|
* positive page count as a precondition. So we use _count directly.
|
|
|
|
*/
|
|
|
|
static void
|
|
|
|
bio_pageinc(struct bio *bio)
|
|
|
|
{
|
|
|
|
struct bio_vec *bv;
|
|
|
|
struct page *page;
|
|
|
|
int i;
|
|
|
|
|
|
|
|
bio_for_each_segment(bv, bio, i) {
|
|
|
|
page = bv->bv_page;
|
|
|
|
/* Non-zero page count for non-head members of
|
|
|
|
* compound pages is no longer allowed by the kernel,
|
|
|
|
* but this has never been seen here.
|
|
|
|
*/
|
|
|
|
if (unlikely(PageCompound(page)))
|
|
|
|
if (compound_trans_head(page) != page) {
|
|
|
|
pr_crit("page tail used for block I/O\n");
|
|
|
|
BUG();
|
|
|
|
}
|
|
|
|
atomic_inc(&page->_count);
|
|
|
|
}
|
|
|
|
}
|
|
|
|
|
|
|
|
static void
|
|
|
|
bio_pagedec(struct bio *bio)
|
|
|
|
{
|
|
|
|
struct bio_vec *bv;
|
|
|
|
int i;
|
|
|
|
|
|
|
|
bio_for_each_segment(bv, bio, i)
|
|
|
|
atomic_dec(&bv->bv_page->_count);
|
|
|
|
}
|
|
|
|
|
|
|
|
static void
|
|
|
|
bufinit(struct buf *buf, struct request *rq, struct bio *bio)
|
|
|
|
{
|
|
|
|
struct bio_vec *bv;
|
|
|
|
|
|
|
|
memset(buf, 0, sizeof(*buf));
|
|
|
|
buf->rq = rq;
|
|
|
|
buf->bio = bio;
|
|
|
|
buf->resid = bio->bi_size;
|
|
|
|
buf->sector = bio->bi_sector;
|
|
|
|
bio_pageinc(bio);
|
|
|
|
buf->bv = bv = &bio->bi_io_vec[bio->bi_idx];
|
|
|
|
buf->bv_resid = bv->bv_len;
|
|
|
|
WARN_ON(buf->bv_resid == 0);
|
|
|
|
}
|
|
|
|
|
|
|
|
static struct buf *
|
|
|
|
nextbuf(struct aoedev *d)
|
|
|
|
{
|
|
|
|
struct request *rq;
|
|
|
|
struct request_queue *q;
|
|
|
|
struct buf *buf;
|
|
|
|
struct bio *bio;
|
|
|
|
|
|
|
|
q = d->blkq;
|
|
|
|
if (q == NULL)
|
|
|
|
return NULL; /* initializing */
|
|
|
|
if (d->ip.buf)
|
|
|
|
return d->ip.buf;
|
|
|
|
rq = d->ip.rq;
|
|
|
|
if (rq == NULL) {
|
|
|
|
rq = blk_peek_request(q);
|
|
|
|
if (rq == NULL)
|
|
|
|
return NULL;
|
|
|
|
blk_start_request(rq);
|
|
|
|
d->ip.rq = rq;
|
|
|
|
d->ip.nxbio = rq->bio;
|
|
|
|
rq->special = (void *) rqbiocnt(rq);
|
|
|
|
}
|
|
|
|
buf = mempool_alloc(d->bufpool, GFP_ATOMIC);
|
|
|
|
if (buf == NULL) {
|
|
|
|
pr_err("aoe: nextbuf: unable to mempool_alloc!\n");
|
|
|
|
return NULL;
|
|
|
|
}
|
|
|
|
bio = d->ip.nxbio;
|
|
|
|
bufinit(buf, rq, bio);
|
|
|
|
bio = bio->bi_next;
|
|
|
|
d->ip.nxbio = bio;
|
|
|
|
if (bio == NULL)
|
|
|
|
d->ip.rq = NULL;
|
|
|
|
return d->ip.buf = buf;
|
2005-04-17 06:20:36 +08:00
|
|
|
}
|
|
|
|
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
/* enters with d->lock held */
|
|
|
|
void
|
|
|
|
aoecmd_work(struct aoedev *d)
|
|
|
|
{
|
|
|
|
if (d->htgt && !sthtith(d))
|
|
|
|
return;
|
2012-10-05 08:16:23 +08:00
|
|
|
while (aoecmd_ata_rw(d))
|
|
|
|
;
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
}
|
|
|
|
|
2006-01-20 02:46:19 +08:00
|
|
|
/* this function performs work that has been deferred until sleeping is OK
|
|
|
|
*/
|
|
|
|
void
|
2006-11-22 22:57:56 +08:00
|
|
|
aoecmd_sleepwork(struct work_struct *work)
|
2006-01-20 02:46:19 +08:00
|
|
|
{
|
2006-11-22 22:57:56 +08:00
|
|
|
struct aoedev *d = container_of(work, struct aoedev, work);
|
2012-10-05 08:16:35 +08:00
|
|
|
struct block_device *bd;
|
|
|
|
u64 ssize;
|
2006-01-20 02:46:19 +08:00
|
|
|
|
|
|
|
if (d->flags & DEVFL_GDALLOC)
|
|
|
|
aoeblk_gdalloc(d);
|
|
|
|
|
|
|
|
if (d->flags & DEVFL_NEWSIZE) {
|
2008-08-25 18:56:07 +08:00
|
|
|
ssize = get_capacity(d->gd);
|
2006-01-20 02:46:19 +08:00
|
|
|
bd = bdget_disk(d->gd, 0);
|
|
|
|
if (bd) {
|
|
|
|
mutex_lock(&bd->bd_inode->i_mutex);
|
|
|
|
i_size_write(bd->bd_inode, (loff_t)ssize<<9);
|
|
|
|
mutex_unlock(&bd->bd_inode->i_mutex);
|
|
|
|
bdput(bd);
|
|
|
|
}
|
2012-10-05 08:16:35 +08:00
|
|
|
spin_lock_irq(&d->lock);
|
2006-01-20 02:46:19 +08:00
|
|
|
d->flags |= DEVFL_UP;
|
|
|
|
d->flags &= ~DEVFL_NEWSIZE;
|
2012-10-05 08:16:35 +08:00
|
|
|
spin_unlock_irq(&d->lock);
|
2006-01-20 02:46:19 +08:00
|
|
|
}
|
|
|
|
}
|
|
|
|
|
2005-04-17 06:20:36 +08:00
|
|
|
static void
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
ataid_complete(struct aoedev *d, struct aoetgt *t, unsigned char *id)
|
2005-04-17 06:20:36 +08:00
|
|
|
{
|
|
|
|
u64 ssize;
|
|
|
|
u16 n;
|
|
|
|
|
|
|
|
/* word 83: command set supported */
|
2008-04-29 16:03:30 +08:00
|
|
|
n = get_unaligned_le16(&id[83 << 1]);
|
2005-04-17 06:20:36 +08:00
|
|
|
|
|
|
|
/* word 86: command set/feature enabled */
|
2008-04-29 16:03:30 +08:00
|
|
|
n |= get_unaligned_le16(&id[86 << 1]);
|
2005-04-17 06:20:36 +08:00
|
|
|
|
|
|
|
if (n & (1<<10)) { /* bit 10: LBA 48 */
|
|
|
|
d->flags |= DEVFL_EXT;
|
|
|
|
|
|
|
|
/* word 100: number lba48 sectors */
|
2008-04-29 16:03:30 +08:00
|
|
|
ssize = get_unaligned_le64(&id[100 << 1]);
|
2005-04-17 06:20:36 +08:00
|
|
|
|
|
|
|
/* set as in ide-disk.c:init_idedisk_capacity */
|
|
|
|
d->geo.cylinders = ssize;
|
|
|
|
d->geo.cylinders /= (255 * 63);
|
|
|
|
d->geo.heads = 255;
|
|
|
|
d->geo.sectors = 63;
|
|
|
|
} else {
|
|
|
|
d->flags &= ~DEVFL_EXT;
|
|
|
|
|
|
|
|
/* number lba28 sectors */
|
2008-04-29 16:03:30 +08:00
|
|
|
ssize = get_unaligned_le32(&id[60 << 1]);
|
2005-04-17 06:20:36 +08:00
|
|
|
|
|
|
|
/* NOTE: obsolete in ATA 6 */
|
2008-04-29 16:03:30 +08:00
|
|
|
d->geo.cylinders = get_unaligned_le16(&id[54 << 1]);
|
|
|
|
d->geo.heads = get_unaligned_le16(&id[55 << 1]);
|
|
|
|
d->geo.sectors = get_unaligned_le16(&id[56 << 1]);
|
2005-04-17 06:20:36 +08:00
|
|
|
}
|
2006-01-20 02:46:19 +08:00
|
|
|
|
|
|
|
if (d->ssize != ssize)
|
2008-02-08 20:20:08 +08:00
|
|
|
printk(KERN_INFO
|
2008-11-25 16:40:37 +08:00
|
|
|
"aoe: %pm e%ld.%d v%04x has %llu sectors\n",
|
|
|
|
t->addr,
|
2006-01-20 02:46:19 +08:00
|
|
|
d->aoemajor, d->aoeminor,
|
|
|
|
d->fw_ver, (long long)ssize);
|
2005-04-17 06:20:36 +08:00
|
|
|
d->ssize = ssize;
|
|
|
|
d->geo.start = 0;
|
2008-02-08 20:20:06 +08:00
|
|
|
if (d->flags & (DEVFL_GDALLOC|DEVFL_NEWSIZE))
|
|
|
|
return;
|
2005-04-17 06:20:36 +08:00
|
|
|
if (d->gd != NULL) {
|
2008-08-25 18:56:07 +08:00
|
|
|
set_capacity(d->gd, ssize);
|
2006-01-20 02:46:19 +08:00
|
|
|
d->flags |= DEVFL_NEWSIZE;
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
} else
|
2006-01-20 02:46:19 +08:00
|
|
|
d->flags |= DEVFL_GDALLOC;
|
2005-04-17 06:20:36 +08:00
|
|
|
schedule_work(&d->work);
|
|
|
|
}
|
|
|
|
|
|
|
|
static void
|
|
|
|
calc_rttavg(struct aoedev *d, int rtt)
|
|
|
|
{
|
|
|
|
register long n;
|
|
|
|
|
|
|
|
n = rtt;
|
2006-09-21 02:36:49 +08:00
|
|
|
if (n < 0) {
|
|
|
|
n = -rtt;
|
|
|
|
if (n < MINTIMER)
|
|
|
|
n = MINTIMER;
|
|
|
|
else if (n > MAXTIMER)
|
|
|
|
n = MAXTIMER;
|
|
|
|
d->mintimer += (n - d->mintimer) >> 1;
|
|
|
|
} else if (n < d->mintimer)
|
|
|
|
n = d->mintimer;
|
2005-04-17 06:20:36 +08:00
|
|
|
else if (n > MAXTIMER)
|
|
|
|
n = MAXTIMER;
|
|
|
|
|
|
|
|
/* g == .25; cf. Congestion Avoidance and Control, Jacobson & Karels; 1988 */
|
|
|
|
n -= d->rttavg;
|
|
|
|
d->rttavg += n >> 2;
|
|
|
|
}
|
|
|
|
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
static struct aoetgt *
|
|
|
|
gettgt(struct aoedev *d, char *addr)
|
|
|
|
{
|
|
|
|
struct aoetgt **t, **e;
|
|
|
|
|
|
|
|
t = d->targets;
|
|
|
|
e = t + NTARGETS;
|
|
|
|
for (; t < e && *t; t++)
|
|
|
|
if (memcmp((*t)->addr, addr, sizeof((*t)->addr)) == 0)
|
|
|
|
return *t;
|
|
|
|
return NULL;
|
|
|
|
}
|
|
|
|
|
2012-10-05 08:16:20 +08:00
|
|
|
static void
|
2012-10-05 08:16:21 +08:00
|
|
|
bvcpy(struct bio_vec *bv, ulong off, struct sk_buff *skb, long cnt)
|
2012-10-05 08:16:20 +08:00
|
|
|
{
|
|
|
|
ulong fcnt;
|
|
|
|
char *p;
|
|
|
|
int soff = 0;
|
|
|
|
loop:
|
|
|
|
fcnt = bv->bv_len - (off - bv->bv_offset);
|
|
|
|
if (fcnt > cnt)
|
|
|
|
fcnt = cnt;
|
|
|
|
p = page_address(bv->bv_page) + off;
|
|
|
|
skb_copy_bits(skb, soff, p, fcnt);
|
|
|
|
soff += fcnt;
|
|
|
|
cnt -= fcnt;
|
|
|
|
if (cnt <= 0)
|
|
|
|
return;
|
|
|
|
bv++;
|
|
|
|
off = bv->bv_offset;
|
|
|
|
goto loop;
|
|
|
|
}
|
|
|
|
|
2012-10-05 08:16:23 +08:00
|
|
|
void
|
|
|
|
aoe_end_request(struct aoedev *d, struct request *rq, int fastfail)
|
|
|
|
{
|
|
|
|
struct bio *bio;
|
|
|
|
int bok;
|
|
|
|
struct request_queue *q;
|
|
|
|
|
|
|
|
q = d->blkq;
|
|
|
|
if (rq == d->ip.rq)
|
|
|
|
d->ip.rq = NULL;
|
|
|
|
do {
|
|
|
|
bio = rq->bio;
|
|
|
|
bok = !fastfail && test_bit(BIO_UPTODATE, &bio->bi_flags);
|
|
|
|
} while (__blk_end_request(rq, bok ? 0 : -EIO, bio->bi_size));
|
|
|
|
|
|
|
|
/* cf. http://lkml.org/lkml/2006/10/31/28 */
|
|
|
|
if (!fastfail)
|
|
|
|
q->request_fn(q);
|
|
|
|
}
|
|
|
|
|
|
|
|
static void
|
|
|
|
aoe_end_buf(struct aoedev *d, struct buf *buf)
|
|
|
|
{
|
|
|
|
struct request *rq;
|
|
|
|
unsigned long n;
|
|
|
|
|
|
|
|
if (buf == d->ip.buf)
|
|
|
|
d->ip.buf = NULL;
|
|
|
|
rq = buf->rq;
|
|
|
|
bio_pagedec(buf->bio);
|
|
|
|
mempool_free(buf, d->bufpool);
|
|
|
|
n = (unsigned long) rq->special;
|
|
|
|
rq->special = (void *) --n;
|
|
|
|
if (n == 0)
|
|
|
|
aoe_end_request(d, rq, 0);
|
|
|
|
}
|
|
|
|
|
2012-10-05 08:16:20 +08:00
|
|
|
static void
|
2012-10-05 08:16:21 +08:00
|
|
|
ktiocomplete(struct frame *f)
|
2012-10-05 08:16:20 +08:00
|
|
|
{
|
2012-10-05 08:16:21 +08:00
|
|
|
struct aoe_hdr *hin, *hout;
|
|
|
|
struct aoe_atahdr *ahin, *ahout;
|
|
|
|
struct buf *buf;
|
|
|
|
struct sk_buff *skb;
|
|
|
|
struct aoetgt *t;
|
|
|
|
struct aoeif *ifp;
|
|
|
|
struct aoedev *d;
|
|
|
|
long n;
|
2012-10-05 08:16:20 +08:00
|
|
|
|
2012-10-05 08:16:21 +08:00
|
|
|
if (f == NULL)
|
2012-10-05 08:16:20 +08:00
|
|
|
return;
|
2012-10-05 08:16:21 +08:00
|
|
|
|
|
|
|
t = f->t;
|
|
|
|
d = t->d;
|
|
|
|
|
|
|
|
hout = (struct aoe_hdr *) skb_mac_header(f->skb);
|
|
|
|
ahout = (struct aoe_atahdr *) (hout+1);
|
|
|
|
buf = f->buf;
|
|
|
|
skb = f->r_skb;
|
|
|
|
if (skb == NULL)
|
|
|
|
goto noskb; /* just fail the buf. */
|
|
|
|
|
|
|
|
hin = (struct aoe_hdr *) skb->data;
|
|
|
|
skb_pull(skb, sizeof(*hin));
|
|
|
|
ahin = (struct aoe_atahdr *) skb->data;
|
|
|
|
skb_pull(skb, sizeof(*ahin));
|
|
|
|
if (ahin->cmdstat & 0xa9) { /* these bits cleared on success */
|
|
|
|
pr_err("aoe: ata error cmd=%2.2Xh stat=%2.2Xh from e%ld.%d\n",
|
|
|
|
ahout->cmdstat, ahin->cmdstat,
|
|
|
|
d->aoemajor, d->aoeminor);
|
|
|
|
noskb: if (buf)
|
2012-10-05 08:16:23 +08:00
|
|
|
clear_bit(BIO_UPTODATE, &buf->bio->bi_flags);
|
2012-10-05 08:16:21 +08:00
|
|
|
goto badrsp;
|
2012-10-05 08:16:20 +08:00
|
|
|
}
|
2012-10-05 08:16:21 +08:00
|
|
|
|
|
|
|
n = ahout->scnt << 9;
|
|
|
|
switch (ahout->cmdstat) {
|
|
|
|
case ATA_CMD_PIO_READ:
|
|
|
|
case ATA_CMD_PIO_READ_EXT:
|
|
|
|
if (skb->len < n) {
|
|
|
|
pr_err("aoe: runt data size in read. skb->len=%d need=%ld\n",
|
|
|
|
skb->len, n);
|
2012-10-05 08:16:23 +08:00
|
|
|
clear_bit(BIO_UPTODATE, &buf->bio->bi_flags);
|
2012-10-05 08:16:21 +08:00
|
|
|
break;
|
|
|
|
}
|
|
|
|
bvcpy(f->bv, f->bv_off, skb, n);
|
|
|
|
case ATA_CMD_PIO_WRITE:
|
|
|
|
case ATA_CMD_PIO_WRITE_EXT:
|
|
|
|
spin_lock_irq(&d->lock);
|
|
|
|
ifp = getif(t, skb->dev);
|
2012-10-05 08:16:27 +08:00
|
|
|
if (ifp)
|
2012-10-05 08:16:21 +08:00
|
|
|
ifp->lost = 0;
|
|
|
|
if (d->htgt == t) /* I'll help myself, thank you. */
|
|
|
|
d->htgt = NULL;
|
|
|
|
spin_unlock_irq(&d->lock);
|
|
|
|
break;
|
|
|
|
case ATA_CMD_ID_ATA:
|
|
|
|
if (skb->len < 512) {
|
|
|
|
pr_info("aoe: runt data size in ataid. skb->len=%d\n",
|
|
|
|
skb->len);
|
|
|
|
break;
|
|
|
|
}
|
|
|
|
if (skb_linearize(skb))
|
|
|
|
break;
|
|
|
|
spin_lock_irq(&d->lock);
|
|
|
|
ataid_complete(d, t, skb->data);
|
|
|
|
spin_unlock_irq(&d->lock);
|
|
|
|
break;
|
|
|
|
default:
|
|
|
|
pr_info("aoe: unrecognized ata command %2.2Xh for %d.%d\n",
|
|
|
|
ahout->cmdstat,
|
|
|
|
be16_to_cpu(get_unaligned(&hin->major)),
|
|
|
|
hin->minor);
|
|
|
|
}
|
|
|
|
badrsp:
|
|
|
|
spin_lock_irq(&d->lock);
|
|
|
|
|
|
|
|
aoe_freetframe(f);
|
|
|
|
|
2012-10-05 08:16:23 +08:00
|
|
|
if (buf && --buf->nframesout == 0 && buf->resid == 0)
|
|
|
|
aoe_end_buf(d, buf);
|
2012-10-05 08:16:21 +08:00
|
|
|
|
2012-10-05 08:16:23 +08:00
|
|
|
aoecmd_work(d);
|
|
|
|
|
|
|
|
spin_unlock_irq(&d->lock);
|
|
|
|
aoedev_put(d);
|
2012-10-05 08:16:21 +08:00
|
|
|
dev_kfree_skb(skb);
|
2012-10-05 08:16:20 +08:00
|
|
|
}
|
|
|
|
|
2012-10-05 08:16:21 +08:00
|
|
|
/* Enters with iocq.lock held.
|
|
|
|
* Returns true iff responses needing processing remain.
|
|
|
|
*/
|
|
|
|
static int
|
|
|
|
ktio(void)
|
|
|
|
{
|
|
|
|
struct frame *f;
|
|
|
|
struct list_head *pos;
|
|
|
|
int i;
|
|
|
|
|
|
|
|
for (i = 0; ; ++i) {
|
|
|
|
if (i == MAXIOC)
|
|
|
|
return 1;
|
|
|
|
if (list_empty(&iocq.head))
|
|
|
|
return 0;
|
|
|
|
pos = iocq.head.next;
|
|
|
|
list_del(pos);
|
|
|
|
spin_unlock_irq(&iocq.lock);
|
|
|
|
f = list_entry(pos, struct frame, head);
|
|
|
|
ktiocomplete(f);
|
|
|
|
spin_lock_irq(&iocq.lock);
|
|
|
|
}
|
|
|
|
}
|
|
|
|
|
|
|
|
static int
|
|
|
|
kthread(void *vp)
|
|
|
|
{
|
|
|
|
struct ktstate *k;
|
|
|
|
DECLARE_WAITQUEUE(wait, current);
|
|
|
|
int more;
|
|
|
|
|
|
|
|
k = vp;
|
|
|
|
current->flags |= PF_NOFREEZE;
|
|
|
|
set_user_nice(current, -10);
|
|
|
|
complete(&k->rendez); /* tell spawner we're running */
|
|
|
|
do {
|
|
|
|
spin_lock_irq(k->lock);
|
|
|
|
more = k->fn();
|
|
|
|
if (!more) {
|
|
|
|
add_wait_queue(k->waitq, &wait);
|
|
|
|
__set_current_state(TASK_INTERRUPTIBLE);
|
|
|
|
}
|
|
|
|
spin_unlock_irq(k->lock);
|
|
|
|
if (!more) {
|
|
|
|
schedule();
|
|
|
|
remove_wait_queue(k->waitq, &wait);
|
|
|
|
} else
|
|
|
|
cond_resched();
|
|
|
|
} while (!kthread_should_stop());
|
|
|
|
complete(&k->rendez); /* tell spawner we're stopping */
|
|
|
|
return 0;
|
|
|
|
}
|
|
|
|
|
2012-10-05 08:16:25 +08:00
|
|
|
void
|
2012-10-05 08:16:21 +08:00
|
|
|
aoe_ktstop(struct ktstate *k)
|
|
|
|
{
|
|
|
|
kthread_stop(k->task);
|
|
|
|
wait_for_completion(&k->rendez);
|
|
|
|
}
|
|
|
|
|
2012-10-05 08:16:25 +08:00
|
|
|
int
|
2012-10-05 08:16:21 +08:00
|
|
|
aoe_ktstart(struct ktstate *k)
|
|
|
|
{
|
|
|
|
struct task_struct *task;
|
|
|
|
|
|
|
|
init_completion(&k->rendez);
|
|
|
|
task = kthread_run(kthread, k, k->name);
|
|
|
|
if (task == NULL || IS_ERR(task))
|
|
|
|
return -ENOMEM;
|
|
|
|
k->task = task;
|
|
|
|
wait_for_completion(&k->rendez); /* allow kthread to start */
|
|
|
|
init_completion(&k->rendez); /* for waiting for exit later */
|
|
|
|
return 0;
|
|
|
|
}
|
|
|
|
|
|
|
|
/* pass it off to kthreads for processing */
|
|
|
|
static void
|
|
|
|
ktcomplete(struct frame *f, struct sk_buff *skb)
|
|
|
|
{
|
|
|
|
ulong flags;
|
|
|
|
|
|
|
|
f->r_skb = skb;
|
|
|
|
spin_lock_irqsave(&iocq.lock, flags);
|
|
|
|
list_add_tail(&f->head, &iocq.head);
|
|
|
|
spin_unlock_irqrestore(&iocq.lock, flags);
|
|
|
|
wake_up(&ktiowq);
|
|
|
|
}
|
|
|
|
|
|
|
|
struct sk_buff *
|
2005-04-17 06:20:36 +08:00
|
|
|
aoecmd_ata_rsp(struct sk_buff *skb)
|
|
|
|
{
|
|
|
|
struct aoedev *d;
|
2012-10-05 08:16:21 +08:00
|
|
|
struct aoe_hdr *h;
|
2005-04-17 06:20:36 +08:00
|
|
|
struct frame *f;
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
struct aoetgt *t;
|
2012-10-05 08:16:21 +08:00
|
|
|
u32 n;
|
2005-04-17 06:20:36 +08:00
|
|
|
ulong flags;
|
|
|
|
char ebuf[128];
|
2005-04-19 13:00:18 +08:00
|
|
|
u16 aoemajor;
|
|
|
|
|
2012-10-05 08:16:21 +08:00
|
|
|
h = (struct aoe_hdr *) skb->data;
|
|
|
|
aoemajor = be16_to_cpu(get_unaligned(&h->major));
|
|
|
|
d = aoedev_by_aoeaddr(aoemajor, h->minor);
|
2005-04-17 06:20:36 +08:00
|
|
|
if (d == NULL) {
|
|
|
|
snprintf(ebuf, sizeof ebuf, "aoecmd_ata_rsp: ata response "
|
|
|
|
"for unknown device %d.%d\n",
|
2012-10-05 08:16:21 +08:00
|
|
|
aoemajor, h->minor);
|
2005-04-17 06:20:36 +08:00
|
|
|
aoechr_error(ebuf);
|
2012-10-05 08:16:21 +08:00
|
|
|
return skb;
|
2005-04-17 06:20:36 +08:00
|
|
|
}
|
|
|
|
|
|
|
|
spin_lock_irqsave(&d->lock, flags);
|
|
|
|
|
2012-10-05 08:16:21 +08:00
|
|
|
n = be32_to_cpu(get_unaligned(&h->tag));
|
2012-10-05 08:16:33 +08:00
|
|
|
f = getframe(d, n);
|
2005-04-17 06:20:36 +08:00
|
|
|
if (f == NULL) {
|
2006-09-21 02:36:49 +08:00
|
|
|
calc_rttavg(d, -tsince(n));
|
2005-04-17 06:20:36 +08:00
|
|
|
spin_unlock_irqrestore(&d->lock, flags);
|
2012-10-05 08:16:23 +08:00
|
|
|
aoedev_put(d);
|
2005-04-17 06:20:36 +08:00
|
|
|
snprintf(ebuf, sizeof ebuf,
|
|
|
|
"%15s e%d.%d tag=%08x@%08lx\n",
|
|
|
|
"unexpected rsp",
|
2012-10-05 08:16:21 +08:00
|
|
|
get_unaligned_be16(&h->major),
|
|
|
|
h->minor,
|
|
|
|
get_unaligned_be32(&h->tag),
|
2005-04-17 06:20:36 +08:00
|
|
|
jiffies);
|
|
|
|
aoechr_error(ebuf);
|
2012-10-05 08:16:21 +08:00
|
|
|
return skb;
|
2005-04-17 06:20:36 +08:00
|
|
|
}
|
2012-10-05 08:16:33 +08:00
|
|
|
t = f->t;
|
2005-04-17 06:20:36 +08:00
|
|
|
calc_rttavg(d, tsince(f->tag));
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
t->nout--;
|
2005-04-17 06:20:36 +08:00
|
|
|
aoecmd_work(d);
|
|
|
|
|
|
|
|
spin_unlock_irqrestore(&d->lock, flags);
|
2012-10-05 08:16:21 +08:00
|
|
|
|
|
|
|
ktcomplete(f, skb);
|
|
|
|
|
|
|
|
/*
|
|
|
|
* Note here that we do not perform an aoedev_put, as we are
|
|
|
|
* leaving this reference for the ktio to release.
|
|
|
|
*/
|
|
|
|
return NULL;
|
2005-04-17 06:20:36 +08:00
|
|
|
}
|
|
|
|
|
|
|
|
void
|
|
|
|
aoecmd_cfg(ushort aoemajor, unsigned char aoeminor)
|
|
|
|
{
|
2008-09-22 13:36:49 +08:00
|
|
|
struct sk_buff_head queue;
|
2005-04-17 06:20:36 +08:00
|
|
|
|
2008-09-22 13:36:49 +08:00
|
|
|
__skb_queue_head_init(&queue);
|
|
|
|
aoecmd_cfg_pkts(aoemajor, aoeminor, &queue);
|
|
|
|
aoenet_xmit(&queue);
|
2005-04-17 06:20:36 +08:00
|
|
|
}
|
|
|
|
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
struct sk_buff *
|
2005-04-17 06:20:36 +08:00
|
|
|
aoecmd_ata_id(struct aoedev *d)
|
|
|
|
{
|
|
|
|
struct aoe_hdr *h;
|
|
|
|
struct aoe_atahdr *ah;
|
|
|
|
struct frame *f;
|
|
|
|
struct sk_buff *skb;
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
struct aoetgt *t;
|
2005-04-17 06:20:36 +08:00
|
|
|
|
2012-10-05 08:16:21 +08:00
|
|
|
f = newframe(d);
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
if (f == NULL)
|
2005-04-17 06:20:36 +08:00
|
|
|
return NULL;
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
|
|
|
|
t = *d->tgt;
|
2005-04-17 06:20:36 +08:00
|
|
|
|
|
|
|
/* initialize the headers & frame */
|
2006-09-21 02:36:49 +08:00
|
|
|
skb = f->skb;
|
2007-10-17 14:27:03 +08:00
|
|
|
h = (struct aoe_hdr *) skb_mac_header(skb);
|
2005-04-17 06:20:36 +08:00
|
|
|
ah = (struct aoe_atahdr *) (h+1);
|
2006-12-22 17:09:21 +08:00
|
|
|
skb_put(skb, sizeof *h + sizeof *ah);
|
|
|
|
memset(h, 0, skb->len);
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
f->tag = aoehdr_atainit(d, t, h);
|
2012-10-05 08:16:21 +08:00
|
|
|
fhash(f);
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
t->nout++;
|
2005-04-17 06:20:36 +08:00
|
|
|
f->waited = 0;
|
|
|
|
|
|
|
|
/* set up ata header */
|
|
|
|
ah->scnt = 1;
|
2009-04-02 03:42:24 +08:00
|
|
|
ah->cmdstat = ATA_CMD_ID_ATA;
|
2005-04-17 06:20:36 +08:00
|
|
|
ah->lba3 = 0xa0;
|
|
|
|
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
skb->dev = t->ifp->nd;
|
2005-04-17 06:20:36 +08:00
|
|
|
|
2006-01-20 02:46:19 +08:00
|
|
|
d->rttavg = MAXTIMER;
|
2005-04-17 06:20:36 +08:00
|
|
|
d->timer.function = rexmit_timer;
|
|
|
|
|
2006-09-21 02:36:49 +08:00
|
|
|
return skb_clone(skb, GFP_ATOMIC);
|
2005-04-17 06:20:36 +08:00
|
|
|
}
|
|
|
|
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
static struct aoetgt *
|
|
|
|
addtgt(struct aoedev *d, char *addr, ulong nframes)
|
|
|
|
{
|
|
|
|
struct aoetgt *t, **tt, **te;
|
|
|
|
|
|
|
|
tt = d->targets;
|
|
|
|
te = tt + NTARGETS;
|
|
|
|
for (; tt < te && *tt; tt++)
|
|
|
|
;
|
|
|
|
|
2008-02-08 20:20:09 +08:00
|
|
|
if (tt == te) {
|
|
|
|
printk(KERN_INFO
|
|
|
|
"aoe: device addtgt failure; too many targets\n");
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
return NULL;
|
2008-02-08 20:20:09 +08:00
|
|
|
}
|
2012-10-05 08:16:21 +08:00
|
|
|
t = kzalloc(sizeof(*t), GFP_ATOMIC);
|
|
|
|
if (!t) {
|
2008-02-08 20:20:09 +08:00
|
|
|
printk(KERN_INFO "aoe: cannot allocate memory to add target\n");
|
aoe: dynamically allocate a capped number of skbs when necessary
What this Patch Does
Even before this recent series of 12 patches to 2.6.22-rc4, the aoe
driver was reusing a small set of skbs that were allocated once and
were only used for outbound AoE commands.
The network layer cannot be allowed to put_page on the data that is
still associated with a bio we haven't returned to the block layer,
so the aoe driver (even before the patch under discussion) is still
the owner of skbs that have been handed to the network layer for
transmission. We need to keep track of these skbs so that we can
free them, but by tracking them, we can also easily re-use them.
The new patch was a response to the behavior of certain network
drivers. We cannot reuse an skb that the network driver still has
in its transmit ring. Network drivers can defer transmit ring
cleanup and then use the state in the skb to determine how many data
segments to clean up in its transmit ring. The tg3 driver is one
driver that behaves in this way.
When the network driver defers cleanup of its transmit ring, the aoe
driver can find itself in a situation where it would like to send an
AoE command, and the AoE target is ready for more work, but the
network driver still has all of the pre-allocated skbs. In that
case, the new patch just calls alloc_skb, as you'd expect.
We don't want to get carried away, though. We try not to do
excessive allocation in the write path, so we cap the number of skbs
we dynamically allocate.
Probably calling it a "dynamic pool" is misleading. We were already
trying to use a small fixed-size set of pre-allocated skbs before
this patch, and this patch just provides a little headroom (with a
ceiling, though) to accomodate network drivers that hang onto skbs,
by allocating when needed. The d->skbpool_hd list of allocated skbs
is necessary so that we can free them later.
We didn't notice the need for this headroom until AoE targets got
fast enough.
Alternatives
If the network layer never did a put_page on the pages in the bio's
we get from the block layer, then it would be possible for us to
hand skbs to the network layer and forget about them, allowing the
network layer to free skbs itself (and thereby calling our own
skb->destructor callback function if we needed that). In that case
we could get rid of the pre-allocated skbs and also the
d->skbpool_hd, instead just calling alloc_skb every time we wanted
to transmit a packet. The slab allocator would effectively maintain
the list of skbs.
Besides a loss of CPU cache locality, the main concern with that
approach the danger that it would increase the likelihood of
deadlock when VM is trying to free pages by writing dirty data from
the page cache through the aoe driver out to persistent storage on
an AoE device. Right now we have a situation where we have
pre-allocation that corresponds to how much we use, which seems
ideal.
Of course, there's still the separate issue of receiving the packets
that tell us that a write has successfully completed on the AoE
target. When memory is low and VM is using AoE to flush dirty data
to free up pages, it would be perfect if there were a way for us to
register a fast callback that could recognize write command
completion responses. But I don't think the current problems with
the receive side of the situation are a justification for
exacerbating the problem on the transmit side.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:05 +08:00
|
|
|
return NULL;
|
|
|
|
}
|
|
|
|
|
2012-10-05 08:16:21 +08:00
|
|
|
d->ntargets++;
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
t->nframes = nframes;
|
2012-10-05 08:16:21 +08:00
|
|
|
t->d = d;
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
memcpy(t->addr, addr, sizeof t->addr);
|
|
|
|
t->ifp = t->ifs;
|
|
|
|
t->maxout = t->nframes;
|
2012-10-05 08:16:21 +08:00
|
|
|
INIT_LIST_HEAD(&t->ffree);
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
return *tt = t;
|
|
|
|
}
|
|
|
|
|
2012-10-05 08:16:27 +08:00
|
|
|
static void
|
|
|
|
setdbcnt(struct aoedev *d)
|
|
|
|
{
|
|
|
|
struct aoetgt **t, **e;
|
|
|
|
int bcnt = 0;
|
|
|
|
|
|
|
|
t = d->targets;
|
|
|
|
e = t + NTARGETS;
|
|
|
|
for (; t < e && *t; t++)
|
|
|
|
if (bcnt == 0 || bcnt > (*t)->minbcnt)
|
|
|
|
bcnt = (*t)->minbcnt;
|
|
|
|
if (bcnt != d->maxbcnt) {
|
|
|
|
d->maxbcnt = bcnt;
|
|
|
|
pr_info("aoe: e%ld.%d: setting %d byte data frames\n",
|
|
|
|
d->aoemajor, d->aoeminor, bcnt);
|
|
|
|
}
|
|
|
|
}
|
|
|
|
|
|
|
|
static void
|
|
|
|
setifbcnt(struct aoetgt *t, struct net_device *nd, int bcnt)
|
|
|
|
{
|
|
|
|
struct aoedev *d;
|
|
|
|
struct aoeif *p, *e;
|
|
|
|
int minbcnt;
|
|
|
|
|
|
|
|
d = t->d;
|
|
|
|
minbcnt = bcnt;
|
|
|
|
p = t->ifs;
|
|
|
|
e = p + NAOEIFS;
|
|
|
|
for (; p < e; p++) {
|
|
|
|
if (p->nd == NULL)
|
|
|
|
break; /* end of the valid interfaces */
|
|
|
|
if (p->nd == nd) {
|
|
|
|
p->bcnt = bcnt; /* we're updating */
|
|
|
|
nd = NULL;
|
|
|
|
} else if (minbcnt > p->bcnt)
|
|
|
|
minbcnt = p->bcnt; /* find the min interface */
|
|
|
|
}
|
|
|
|
if (nd) {
|
|
|
|
if (p == e) {
|
|
|
|
pr_err("aoe: device setifbcnt failure; too many interfaces.\n");
|
|
|
|
return;
|
|
|
|
}
|
2012-10-05 08:16:34 +08:00
|
|
|
dev_hold(nd);
|
2012-10-05 08:16:27 +08:00
|
|
|
p->nd = nd;
|
|
|
|
p->bcnt = bcnt;
|
|
|
|
}
|
|
|
|
t->minbcnt = minbcnt;
|
|
|
|
setdbcnt(d);
|
|
|
|
}
|
|
|
|
|
2005-04-17 06:20:36 +08:00
|
|
|
void
|
|
|
|
aoecmd_cfg_rsp(struct sk_buff *skb)
|
|
|
|
{
|
|
|
|
struct aoedev *d;
|
|
|
|
struct aoe_hdr *h;
|
|
|
|
struct aoe_cfghdr *ch;
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
struct aoetgt *t;
|
2005-04-19 13:00:20 +08:00
|
|
|
ulong flags, sysminor, aoemajor;
|
2005-04-17 06:20:36 +08:00
|
|
|
struct sk_buff *sl;
|
2012-10-05 08:16:23 +08:00
|
|
|
struct sk_buff_head queue;
|
2006-09-21 02:36:49 +08:00
|
|
|
u16 n;
|
2005-04-17 06:20:36 +08:00
|
|
|
|
2012-10-05 08:16:23 +08:00
|
|
|
sl = NULL;
|
2007-10-17 14:27:03 +08:00
|
|
|
h = (struct aoe_hdr *) skb_mac_header(skb);
|
2005-04-17 06:20:36 +08:00
|
|
|
ch = (struct aoe_cfghdr *) (h+1);
|
|
|
|
|
|
|
|
/*
|
|
|
|
* Enough people have their dip switches set backwards to
|
|
|
|
* warrant a loud message for this special case.
|
|
|
|
*/
|
2008-07-04 15:28:32 +08:00
|
|
|
aoemajor = get_unaligned_be16(&h->major);
|
2005-04-17 06:20:36 +08:00
|
|
|
if (aoemajor == 0xfff) {
|
2006-09-21 02:36:51 +08:00
|
|
|
printk(KERN_ERR "aoe: Warning: shelf address is all ones. "
|
2006-09-21 02:36:49 +08:00
|
|
|
"Check shelf dip switches.\n");
|
2005-04-17 06:20:36 +08:00
|
|
|
return;
|
|
|
|
}
|
2012-10-05 08:16:32 +08:00
|
|
|
if (h->minor >= NPERSHELF) {
|
|
|
|
pr_err("aoe: e%ld.%d %s, %d\n",
|
|
|
|
aoemajor, h->minor,
|
|
|
|
"slot number larger than the maximum",
|
|
|
|
NPERSHELF-1);
|
|
|
|
return;
|
|
|
|
}
|
2005-04-17 06:20:36 +08:00
|
|
|
|
|
|
|
sysminor = SYSMINOR(aoemajor, h->minor);
|
2005-04-19 13:00:17 +08:00
|
|
|
if (sysminor * AOE_PARTITIONS + AOE_PARTITIONS > MINORMASK) {
|
2006-09-21 02:36:51 +08:00
|
|
|
printk(KERN_INFO "aoe: e%ld.%d: minor number too large\n",
|
2005-04-19 13:00:17 +08:00
|
|
|
aoemajor, (int) h->minor);
|
2005-04-17 06:20:36 +08:00
|
|
|
return;
|
|
|
|
}
|
|
|
|
|
2006-09-21 02:36:49 +08:00
|
|
|
n = be16_to_cpu(ch->bufcnt);
|
2008-02-08 20:20:07 +08:00
|
|
|
if (n > aoe_maxout) /* keep it reasonable */
|
|
|
|
n = aoe_maxout;
|
2005-04-17 06:20:36 +08:00
|
|
|
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
d = aoedev_by_sysminor_m(sysminor);
|
2005-04-17 06:20:36 +08:00
|
|
|
if (d == NULL) {
|
2006-09-21 02:36:51 +08:00
|
|
|
printk(KERN_INFO "aoe: device sysminor_m failure\n");
|
2005-04-17 06:20:36 +08:00
|
|
|
return;
|
|
|
|
}
|
|
|
|
|
|
|
|
spin_lock_irqsave(&d->lock, flags);
|
|
|
|
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
t = gettgt(d, h->src);
|
|
|
|
if (!t) {
|
|
|
|
t = addtgt(d, h->src, n);
|
2012-10-05 08:16:23 +08:00
|
|
|
if (!t)
|
|
|
|
goto bail;
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
}
|
2012-10-05 08:16:27 +08:00
|
|
|
n = skb->dev->mtu;
|
|
|
|
n -= sizeof(struct aoe_hdr) + sizeof(struct aoe_atahdr);
|
|
|
|
n /= 512;
|
|
|
|
if (n > ch->scnt)
|
|
|
|
n = ch->scnt;
|
|
|
|
n = n ? n * 512 : DEFAULTBCNT;
|
|
|
|
setifbcnt(t, skb->dev, n);
|
2006-01-20 02:46:19 +08:00
|
|
|
|
|
|
|
/* don't change users' perspective */
|
2012-10-05 08:16:23 +08:00
|
|
|
if (d->nopen == 0) {
|
|
|
|
d->fw_ver = be16_to_cpu(ch->fwver);
|
|
|
|
sl = aoecmd_ata_id(d);
|
2005-04-17 06:20:36 +08:00
|
|
|
}
|
2012-10-05 08:16:23 +08:00
|
|
|
bail:
|
2005-04-17 06:20:36 +08:00
|
|
|
spin_unlock_irqrestore(&d->lock, flags);
|
2012-10-05 08:16:23 +08:00
|
|
|
aoedev_put(d);
|
2008-09-22 13:36:49 +08:00
|
|
|
if (sl) {
|
|
|
|
__skb_queue_head_init(&queue);
|
|
|
|
__skb_queue_tail(&queue, sl);
|
|
|
|
aoenet_xmit(&queue);
|
|
|
|
}
|
2005-04-17 06:20:36 +08:00
|
|
|
}
|
|
|
|
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
void
|
|
|
|
aoecmd_cleanslate(struct aoedev *d)
|
|
|
|
{
|
|
|
|
struct aoetgt **t, **te;
|
|
|
|
|
|
|
|
d->mintimer = MINTIMER;
|
2012-10-05 08:16:27 +08:00
|
|
|
d->maxbcnt = 0;
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
|
|
|
|
t = d->targets;
|
|
|
|
te = t + NTARGETS;
|
2012-10-05 08:16:27 +08:00
|
|
|
for (; t < te && *t; t++)
|
aoe: handle multiple network paths to AoE device
A remote AoE device is something can process ATA commands and is identified by
an AoE shelf number and an AoE slot number. Such a device might have more
than one network interface, and it might be reachable by more than one local
network interface. This patch tracks the available network paths available to
each AoE device, allowing them to be used more efficiently.
Andrew Morton asked about the call to msleep_interruptible in the revalidate
function. Yes, if a signal is pending, then msleep_interruptible will not
return 0. That means we will not loop but will call aoenet_xmit with a NULL
skb, which is a noop. If the system is too low on memory or the aoe driver is
too low on frames, then the user can hit control-C to interrupt the attempt to
do a revalidate. I have added a comment to the code summarizing that.
Andrew Morton asked whether the allocation performed inside addtgt could use a
more relaxed allocation like GFP_KERNEL, but addtgt is called when the aoedev
lock has been locked with spin_lock_irqsave. It would be nice to allocate the
memory under fewer restrictions, but targets are only added when the device is
being discovered, and if the target can't be added right now, we can try again
in a minute when then next AoE config query broadcast goes out.
Andrew Morton pointed out that the "too many targets" message could be printed
for failing GFP_ATOMIC allocations. The last patch in this series makes the
messages more specific.
Signed-off-by: Ed L. Cashin <ecashin@coraid.com>
Cc: Greg KH <greg@kroah.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-08 20:20:00 +08:00
|
|
|
(*t)->maxout = (*t)->nframes;
|
|
|
|
}
|
2012-10-05 08:16:21 +08:00
|
|
|
|
2012-10-05 08:16:23 +08:00
|
|
|
void
|
|
|
|
aoe_failbuf(struct aoedev *d, struct buf *buf)
|
|
|
|
{
|
|
|
|
if (buf == NULL)
|
|
|
|
return;
|
|
|
|
buf->resid = 0;
|
|
|
|
clear_bit(BIO_UPTODATE, &buf->bio->bi_flags);
|
|
|
|
if (buf->nframesout == 0)
|
|
|
|
aoe_end_buf(d, buf);
|
|
|
|
}
|
|
|
|
|
|
|
|
void
|
|
|
|
aoe_flush_iocq(void)
|
2012-10-05 08:16:21 +08:00
|
|
|
{
|
|
|
|
struct frame *f;
|
|
|
|
struct aoedev *d;
|
|
|
|
LIST_HEAD(flist);
|
|
|
|
struct list_head *pos;
|
|
|
|
struct sk_buff *skb;
|
|
|
|
ulong flags;
|
|
|
|
|
|
|
|
spin_lock_irqsave(&iocq.lock, flags);
|
|
|
|
list_splice_init(&iocq.head, &flist);
|
|
|
|
spin_unlock_irqrestore(&iocq.lock, flags);
|
|
|
|
while (!list_empty(&flist)) {
|
|
|
|
pos = flist.next;
|
|
|
|
list_del(pos);
|
|
|
|
f = list_entry(pos, struct frame, head);
|
|
|
|
d = f->t->d;
|
|
|
|
skb = f->r_skb;
|
|
|
|
spin_lock_irqsave(&d->lock, flags);
|
|
|
|
if (f->buf) {
|
|
|
|
f->buf->nframesout--;
|
|
|
|
aoe_failbuf(d, f->buf);
|
|
|
|
}
|
|
|
|
aoe_freetframe(f);
|
|
|
|
spin_unlock_irqrestore(&d->lock, flags);
|
|
|
|
dev_kfree_skb(skb);
|
2012-10-05 08:16:23 +08:00
|
|
|
aoedev_put(d);
|
2012-10-05 08:16:21 +08:00
|
|
|
}
|
|
|
|
}
|
|
|
|
|
|
|
|
int __init
|
|
|
|
aoecmd_init(void)
|
|
|
|
{
|
|
|
|
INIT_LIST_HEAD(&iocq.head);
|
|
|
|
spin_lock_init(&iocq.lock);
|
|
|
|
init_waitqueue_head(&ktiowq);
|
|
|
|
kts.name = "aoe_ktio";
|
|
|
|
kts.fn = ktio;
|
|
|
|
kts.waitq = &ktiowq;
|
|
|
|
kts.lock = &iocq.lock;
|
|
|
|
return aoe_ktstart(&kts);
|
|
|
|
}
|
|
|
|
|
|
|
|
void
|
|
|
|
aoecmd_exit(void)
|
|
|
|
{
|
|
|
|
aoe_ktstop(&kts);
|
2012-10-05 08:16:23 +08:00
|
|
|
aoe_flush_iocq();
|
2012-10-05 08:16:21 +08:00
|
|
|
}
|