Tuesday, November 23, 2010

CAPWAP Split-MAC Architecture Overview

One of the key principles behind the LWAPP and CAPWAP protocol architecture is the notion of a split 802.11 media access control. Since the real processing power and smart feature set of the architecture is implemented in controllers, some functions need to be performed in the controller instead of the access point. This concept is called "Split-MAC" by Cisco and most other controller-based vendors.

The AP and controller are linked by the CAPWAP protocol using both a "control" channel for access point management, configuration, and control, and a "data" channel for forwarding of user traffic between the two entities in the cases where user traffic is tunneled all the way to the controller (central bridging). These two channels are nothing more than CAPWAP encapsulated UDP packets using port 5246 (control) and 5247 (data) since Cisco code version 5.2. Earlier versions of code used the LWAPP protocol, which was CAPWAP's predecessor, and use UDP ports 12223 (control) and 12222 (data).

It is important for wireless engineers designing, deploying, administering, and troubleshooting solutions using this type of architecture to understand the functions carried out by the controller versus the access point.

The industry is currently in a transition back to a de-centralized model, with local data bridging coming into higher demand as 802.11n data rates strain controller bandwidth capacity and branch offices struggle to cost-justify the additional expense of controllers. This is evident with the emergence of Cisco H-REAP, Aruba RAP, Motorola Adaptive APs, and taken to the extreme by Aerohive in their controller-less architecture. This trend will only continue, but engineers will still be required to fully understand the split-MAC concept even under these circumstances as the large vendors are likely to require centralized controllers for some control-plane functions.

The split-MAC functionality is divided between controller and AP in the following fashion:

Controller Responsibilities:

  • Security management (policy enforcement, rogue detection, etc.)
  • Configuration and firmware management
  • Northbound management interfaces
  • Non real-time 802.11 MAC functions
    • Association, Dis-Association, Re-Association
    • 802.11e/WMM Resource Reservation (CAC, TSPEC, etc.)
    • 802.1x/EAP Authentication
    • Encryption Key Management
  • 802.11 Distribution Services
  • Wired and Wireless Integration Services

Access Point Responsibilities:

  • Real-Time 802.11 MAC Functions
    • Beacon generation
    • Probe responses
    • Informs WLC of client probe requests
    • Power management and packet buffering
    • 802.11e/WMM scheduling and queuing
    • MAC layer data encryption and decryption
    • 802.11 control messages (ACK, RTS/CTS)
  • Data encapsulation and de-capsulation via CAPWAP
  • Fragmentation and re-assembly
  • RF spectral analysis
  • WLAN IDS signature analysis

In future posts, I detail how CAPWAP APs discover, select, join, and maintain association with a controller.

Cheers,
Andrew

Wednesday, November 17, 2010

Cisco WLC QoS Profile Bugs

There is a pretty devious little bug ID in the Cisco WLC QoS profile settings that can lead to problematic traffic forwarding for users and a pretty severe disruption to user experience.

Recently, during a guest wireless deployment, our wireless engineering team designed a solution to minimize potential impact of guest Wi-Fi devices on our production network by implementing QoS bandwidth contracts to guest users. This is accomplished through the Wireless > QoS > Profiles section of the wireless controller configuration, as shown below.


When this QoS profile is applied to a WLAN, each user's bandwidth is limited to the specified average and burst data rates (in Kbps). Additionally, the wired QoS protocol tagging feature uses 802.1p for layer 2 class of service integration with the wired network switches.

Once implemented, we began experiencing random TCP sessions dropping as well as random UDP packet loss. Scratching our heads, we started digging into what changed recently. Besides upgrading to 7.0.98.0 code, the only other change was enforcement of QoS bandwidth contracts using the Bronze QoS profile applied to the guest wireless WLAN. It also happened that the guest user base were the only individuals reporting problems. They were having IPSec and SSL VPN sessions dropping numerous times throughout each day. Definitely not a productive environment to support many corporate business partners whom we rely on to help us get work done.

Many packet traces at *multiple* points throughout the network later, we found return packets to clients missing, and apparently being dropped by the local controller. We expected that tracing packets in both CAPWAP and Mobility Ether-IP tunnels would be fairly complex, but were pleasantly surprised to see that Wireshark had protocol dissectors for both protocols! Yeah Wireshark!

Also, having a network of remote sniffers deployed throughout various points in your network is a godsend for remote configuration and packet capture capabilities. Thank you to our internal performance services team for having this capability! I can only imagine how painful it would have been to have to physically visit each point in the network to setup a SPAN port or other manual network sniffer solution.

So, knowing the local controller (as opposed to the DMZ anchor, or any other network equipment) was dropping the packets, it was a fairly simple matter of figuring out what changed to cause the issue. Rather than downgrading code, we first reverted the configuration to it's state prior to the problems by removing the QoS bandwidth contracts. Voila, problem magically disappeared.

Searching the Cisco documentation of this feature revealed no clues or warnings around best practice use of this feature (normally, if Cisco does not want a customer to modify a default value without contacting TAC or Advanced Services they will explicitly state that in the configuration guide). No warnings, okay - what is going on here, why is this feature not working correctly. Hmmm.... BUG?!

How about that, numerous bugs exist for the QoS profile settings for 802.1p tagging and bandwidth contracts. Here's one interesting one:

CSCsz20162 - WLC5500: QoS rate limiting feature is not accurate
This bug was resolved in 7.0.98.0 code, or was it?

Additionally, here are some other bugs not fixed until code versions after 7.0.98.0 (all fixes apply to a later release of code):

CSCth94887 - 5500: 802.1p markings not working for pings, EoIP and FastCache mode
CSCth90962 - Document QoS 802.1p tagging blocks traffic on untagged interfaces
CSCte64638 - Clients cannot ping or get IP address with 802.1p QoS profile
CSCti62070 - Per-User bandwidth contract blocks all traffic when set to 0
CSCte53175 - Per-User bandwidth contract blocks all traffic when set to 0

None of these directly relate to our issue, but there are enough bugs on the feature to make me think I may have found a new one. Additionally, since our issue cleared right up when removing the bandwidth contracts, a bug seems to be the likely bet here.

I also find it interesting that Cisco's Bug Toolkit does not include lookups for their 5508 series wireless controllers. So I'm stuck looking up bugs for the 4400 series, hoping that everything affecting 5508's also affects 4400's.

As Ethan Banks posted over in his blog, engineers are constantly opening support cases with vendors for bugs. "All the bugs. All the time, bugs. Bugs and bugs and bugs. Buggy bug bugs. AHHH!" I can relate to that.


So, what have we learned from this experience?
  1. Document all changes to systems. There is a reason that change management practices and processes exist in most organizations. Use it, live by it, it will save your butt.
  2. Test your changes prior to production implementation. We did test this change and still failed to catch the issue. But we will update our testing procedures. We catch most issues with testing, but some still fall through the cracks. Learn from them and update your testing plans accordingly.
  3. Prepare troubleshooting tools ahead of time. Without remote packet capture capability, it might have taken us 10x longer to figure out where the issue was occurring. We had new addressing space, routing, firewalls, proxies, etc. all involved that I didn't discuss for brevity, but it could have been anywhere. Take a pro-active approach to deploying troubleshooting and monitoring tools. This way, when a problem comes up you can react quickly and execute on an emergency response plan. They do this for fire drills and emergency responders, take the time to do it for your network!
  4. Admit mistakes, take responsibility, and fix the issue. I've seen too many individuals in IT attempt to deny issues and gloss over relevant information for fear of looking bad, incompetent, or just plain "not perfect" at what they do. Some even go so far as to lie about log data (or selectively focus on data that reinforces their position while discounting data that does not). This leads to longer issue resolution times, worse business impact, and once the root cause is determined, which it will be, they end up losing credibility and trust of everyone involved. Instead, present all data that could possibly be involved in the issue up front, and if it is your issue take responsibility. Everyone makes mistakes or simply cannot do a *perfect* job. Even the best engineers, including CCIEs (yes, you're fallible too)!
For the time being, I would recommend avoiding per-user bandwidth contracts in your wireless LAN controller QoS profiles. 


Best of luck out there with your bugs ;)
Andrew

Monday, November 15, 2010

Nuts About Nets AirHORN Overview

AirHORN is an RF signal generator capable of transmitting standard Wi-Fi modulated signals on the 2.4 GHz and 5 GHz ISM and UNII bands.

*Note – AirHORN does not support signal transmission in the 5GHz UNII-3 and ISM band (5.725 – 5.875 MHz).

The dual-band version of the product comprises a single USB adapter with internal antenna and a USB mount with cable extender for optimal device orientation and polarization. The single-band version comes with an external RP-SMA connector and 5dBi omni-directional antenna, offering the flexibility to use alternate external antennas. Installation of the product requires an available USB 2.0 port and Microsoft .NET 2.0 Framework. Older laptops and workstations with USB 1.1 ports will not work properly and the software will not initialize.


AirHORN is useful for RF antenna engineers when researching and developing wireless antenna propagation, amplifier performance, and receiver operation to ensure accurate signal transmission and reception according to desired specifications. It can also be used post-manufacture by wireless LAN engineers when designing and installing Wi-Fi networks, and is especially useful for directional and semi-directional antenna alignment. It can also be used in the absence of a Wi-Fi access point during pre-installation site surveys to identify optimal AP placement for complete RF coverage assessment as well as to avoid RF dead spots.

The signal generated by AirHORN is viewable with any spectrum analysis tool, such as the Cisco Spectrum Expert as shown in this demonstration.


Another good use for AirHORN is to combine its use with other network performance evaluation tools, such as NetStress, IxChariot, or Iperf to determine the negative impact that co-channel interference (CCI) and adjacent channel interference (ACI) can have on a Wi-Fi network from other RF sources. When combined, these tools can be used as a valuable training classroom lab aid to demonstrate these concepts to inexperienced wireless LAN engineers, to identify and measure the effects of neighboring RF impact to a Wi-Fi network, to generate impact assessment reports for management, and develop internal best practices around AP placement and channel overlap to minimize negative impact.

The product can also function as a denial of service (DoS) tool to cause severe disruption to a wireless LAN network since it utilizes 100% of available airtime (duty cycle). The signal generated by AirHORN is capable of completely wiping out a Wi-Fi network. However, the product is specifically designed to comply with all FCC regulations and IEEE standards for power output and transmission power is limited to 17dBm (50 mW). Also, AirHORN causes performance degradation in part because the signal transmission does not adhere to IEEE 802.11 medium contention rules and acts as a continuous transmitter.

Wireless LAN engineers should know that the tool is not automatically classified by Cisco CleanAir spectrum analysis access points. This is due to the CleanAir system architecture which splits RF signal processing between the Wi-Fi chipset and the SAgE spectrum chipset, with the focus of the SAgE chipset on non-Wi-Fi interference classification. The Wi-Fi modulated AirHORN RF signal is sent to the Wi-Fi chipset and are not processed by the SAgE chipset, with the resulting energy being interpreted as Wi-Fi adjacent channel interference and contributing to overall channel utilization. Therefore, the SAgE chipset and CleanAir system cannot classify the signal.

Product Link:

Cheers,
Andrew