Wednesday, May 22, 2019

VOIP down. Data up. Initial look doesn't reveal anything obvious. No redundancy. Reload equipment or keep data up while isolating the problem? Why or why not?

I'm thinking more on the side of branch offices where we don't have high availability equipment. Our job is all about uptime. But sometimes power cycling a couple of suspect devices is faster than finding the exact problem. Would you take down another service in hopes to fix another one more quickly?

edit: what if it were data down and VOIP up?



Having issues connecting Fortigate 60E to Comcast Metro Fiber.

We have a super simple set-up, but I just cannot get this to work. Our network only has outbound traffic, nothing coming in for RDP, web hosting, or email since everything is hosted off site for security reasons. Because of this we don't have any WAN block IPs from Comcast except the one used for our Fortigate 60e, or whatever device is attached to the Ciena switch for the fiber.

The configuration Comcast gave me is as follows:
Link IP Address: 50.XXX.XXX.208/30
Gateway: 50.XXX.XXX.209
Layer 3 IP: 50.XXX.XXX.210
Layer 3 Subnet Mask: 255.255.255.252

I've configured the WAN port on the FG to be set to the Layer 3 IP, configured the IPv4 policy, and get nothing. If I swap from the fiber installation to the cable modem that is still active, and reconfigure the IP for that, the network has internet. I can ping the Ciena switch and the 50.XXX.XXX.210 ip for the WAN port from our server or any of our terminals, but still no internet.
So I plugged a laptop directly into the Ciena switch, set it's IP as what the WAN port was configured to, and everything works.
Out of desperation I've even tried configuring static route for the Link IP address and gateway, but still don't get anything.
I feel like I'm missing something obvious that I need to enable, but I can't for the life of me remember what it is.
Any advice?



Tunnel From Cisco ASAv to Palo Alto

Hi guys, I've got an ASAv sitting in Azure. Let's say it has an "outside" interface with ip address of 192.168.1.1 for the Azure network.

In Azure, a static Public IP Address is assigned to that interface. We'll call it 10.10.10.2 (yes this is a private range, just an example).

When IKEv1 tries to negotiate Phase 1, it fails: IKE phase-1 negotiation is failed. Peer's ID payload 192.168.1.1 (type ipaddr) does not match a configured IKE gateway.

Now obviously, my IKE gateway is specifying the public Azure IP address of the ASAv... but when it gets the packet from the ASAv, the payload says 192.168.1.1 because that's what the ASA thinks its own IP address is.

I've got many tunnels to physical ASAs that don't seem to have this problem. I've been researching for a couple hours, but don't understand how I can resolve this, or why it doesn't happen on the other ASAs.

I'm a Server guy by trade, so maybe I'm missing something obvious here?



Interoperable QOS Woes

Just when I think I've gotten a handle on QOS something knocks me down.

I'm currently in an MPLS Environment where we need to work within a 4 queue structure with our carrier. We do get to choose various algorithms from them and queuing structures.

We were running zero qos before but predictably when we got congestion tons of user reports about critical traffic being dropped and the support desks response was to hunt down everyone browsing youtube on their breaks and shut them down... This wasn't sustainable.

So I put in a decent amount of time trying to finally learn a command other than auto-qos and we put in a hierarchical policy that I thought looked pretty good:

Example Parent Policy

policy-map 200MB_SHAPE_CC_EDGE_WAN class class-default shape average 200000000 service-policy CC_EDGE_TO_MPLS 

Example Child Policy

policy-map CC_EDGE_TO_MPLS class VOIP priority percent 30 class NCONTROL priority percent 5 class CRITICAL bandwidth remaining percent 60 random-detect dscp-based random-detect ecn class class-default random-detect dscp-based random-detect ecn fair-queue 

We chose this method because it allows us to have parent policies for all our different bandwidth metrics. Our largest site is 700mb and our smallest is 50mb.

After we implemented this all was right with the world for the last six months. We no longer get user reports when we max out our bandwidth.

Until this week. This week one of our larger sites at 200mb finally started hitting their max and our telecom team came running over showing that the RTP streams are experiencing loss.

Weird I thought... check the VOIP buffer and no drops. I reach out to our carrier and surprisingly they are very helpful they point that the issue may be with some of the QOS settings that aren't viable for a 200 MB link.

Specifically they stated the following three items:

rate correctly set at 200M but the bc and be are over scaled at 800000 bits each equaling at 100KB burst... the tolerence is 64kb... recommend adjusting the BC and BE values manually to define at 512000 each if the CPE will allow

Raise the queue limit of 833 to at least 1000 as it is too small for the 200M service.

As for the nested QOS, it is not an exact match to the ordered network of 30-06-42-22. The output shows the cpe is set 30-05- bandwidth remaining 60%. This entials that your AF tagging is not allocating 42% of your CIR, rather your AF and BE combined are claiming the remaining 60% of the bandwidth which could account for further drops.

So I have a few concerns/questions I have been trying to google an understand but it seems like every recommendation points me in different directions. Hopefully someone here who has much more experience can weigh in and give me a hand.

Platform is ASR1001

On the ASR when going to set the BC and BE values the context sensitive help literally recommends against setting the BC and BE manually saying an algo will find the best value. Is this safe to ignore? Is there another Cisco feature that I should be using in order to more safely scale this correctly to match the carrier?

Queue Limits - I don't seem to have any control over the values that are set for the priority queues or the parent shaper. Is there something else I can do here?

The nested QOS is a fair issue. Cisco allows two priority queues and everything I found suggested VOIP and network control traffic should go in those priority queues. Our carrier only has one priority queue for EF traffic only. CS7/6 traffic would go into their P2 queue.

Is it better to just adjust our network control out of priority so it's easier to match the carrier? How do you all handle differing carrier policies and queues?



Why would you NOT want to let higher QoS/CoS tunnels expand?

I'm looking at an implementation where lower QoS/CoS tunnels are allowed to expand if there's unallocated bandwidth, but higher QoS/CoS tunnels are not allowed to. I'm having a hard time thinking of a practical reason for this that you would see in the wild.

The only thing I can come up with is that someone could spoof high priority traffic and cause a DoS attack, but in many cases you shouldn't be accepting the tagging of the incoming traffic, you should be deciding for yourself how you'll tag it.

Are there any other commonly seen reasons someone would want it this way?



Shared firewall with multiple customers

Hi everyone,

We run a small datacenter and mostly everything is just L2 and each customer having their own firewall. We do run a shared firewall on a Sophos SG210 running UTM, and each customer having their own VLAN, and we assign them a public ipv4 address which we just NAT to their specfic VLAN, and we not that happy with that solution. So now, we're looking for a firewall that is meant to be used for multiple customers. What kind of firewall would guys suggest as a multitenant firewall?

Thanks!



LPT: A shitty laptop and dumpcap for intermittent issues on a budget.

It happens to all of us, some weird random problem that happens after-hours or some especially whiny end user. It'd be a hell of a lot easier if you had a historical capture of the data within that timeframe right? Well, if you have a shitty desktop or laptop with a non-flash based HDD (more room typically) you can make that happen.

1.SPAN, RSPAN or ERSPAN (or a hub but that's a bad idea long term) the port or traffic to your laptop using the Googles (you want a port in the path of the affected user or their port)

https://ccie-or-null.net/2011/04/04/configure-span-session/

  1. Setup dumpcap

https://www.youtube.com/watch?v=WJM9wSR8PVM

  1. Review those sweet, sweet PCAPs around the timeframe and begin to correlate what's happening in your infrastructure around that timeframe.

4: ???

5: Profit

Edit: Add a second NIC to be able to manage the box, or you won't be able to get to it as a SPAN destination.



ISP Gateway providing ARP replies for our IPspace with differing VRRP Mac addresses?

Greetings Everyone!

Apologies for the length of the post, as I'm trying to provide as much context and documentation as I can. Trying to wrap my head around an issue we're having here. We have dual firewalls in a HA failover config. each Firewall has a physical IP and several Virtual IPs configured for High Availability VRRP when the firewalls are failed over. the VRRP Virtual Mac addresses all start with 00:00:5e:00:01:0-VHID. Depending on the response, there are stretches of time where we lose connectivity - somtimes after a couple of hours, sometimes after a few days ; I'm convinced as a result of some sort of arp cache issues. The MAC addresses can be traced to our systems, so it's not a matter of dupicate IPs in their extended network since we share a subnet with their other customers. I found no rogue or unidentifiable Mac addresses, it's just that that sometimes THE ISP gateway responds with the Physical Interface MAC and sometimes with the VRRP Virtual MAC. We maintain 2 other HA Firewall Configs with differing IPspace and ISPs that have the same type of configs, both of those have been trouble free. This ISP is the only one that has been causing an issue, and it seems more frequent as time goes on.

I welcome any insight at this point. I feel like I'm going insane.

My Question:

  • Their gateway literally responds like a know-it-all grammar school kid to EVERY single arp request, answering on behalf of both of our physical IPs and virtual IPs. Sometimes the MAC addresses are the same, sometimes they differ. Sometimes it advertises the Physical MAC address and sometimes it advertises the VRRP Virtual MAC address (in the case of the Virtual IPs). For each ARP request, I receive 2 ARP replies, 1 from our system that's the target of the arping, and 1 from their gateway. Is this normal behavior that should be expected? Literally it answers for EVERYTHING associated with our IPspace.

As an example, here's the ARP table on our standby firewall with the primary firewall as active. I sent the below output (as well as traceroutes etc) to their tech team, and I got the verbal equivalent of eyes glazing over. Their level 3 support defaulted to "reboot the modem" which is something we've done at least once a week each time anyway. Rebooting the modem brought a couple of the Virtual IPs back, others remain an issue as noted below.

(ISP_GATEWAY) at (ISP_GW_MAC) on bge4 expires in 1173 seconds [ethernet]

(VIRTUAL_IP1) at (FW1_PHYS_MAC) on bge4 expires in 388 seconds [ethernet]

(PHYSICAL_FW2) at (FW2_PHYS_MAC) on bge4 permanent [ethernet]

(VIRTUAL_IP2) at (VIP2_VIRT_MAC) on bge4 expires in 1193 seconds [ethernet]

(VIRTUAL_IP3) at (VIP3_VIRTUAL_MAC) on bge4 expires in 1183 seconds [ethernet]

(VIRTUAL_IP4) at (FW1_PHYS_MAC) on bge4 expires in 1170 seconds [ethernet]

  • arping output and associated tcpdump. In the below case, it's responding with the physical MAC address of our primary firewall, while the primary firewall is responding with the Virtual VRRP MAC.

ARPING VIRT_IP_3

60 bytes from FW1_PHYS_MAC (VIRT_IP_3): index=0 time=165.729 usec

60 bytes from ISP_GW_MAC (VIRT_IP_3): index=1 time=10.383 msec

60 bytes from FW1_PHYS_MAC (VIRT_IP_3): index=2 time=183.767 usec

60 bytes from ISP_GW_MAC (VIRT_IP_3): index=3 time=12.337 msec

60 bytes from FW1_PHYS_MAC (VIRT_IP_3): index=4 time=181.841 usec

60 bytes from ISP_GW_MAC (VIRT_IP_3): index=5 time=104.296 msec

10:35:53.667220 ARP, Request who-has VIRT_IP_3 tell FW2_PHYS_IP, length 44

10:35:53.667385 ARP, Reply VIRT_IP_3 is-at VIP3_VIRT_MAC (oui IANA), length 46

10:35:53.677589 ARP, Reply VIRT_IP_3 is-at FW1_PHYS_MAC (oui Unknown), length 46

10:35:54.667351 ARP, Request who-has VIRT_IP_3 tell FW2_PHYS_IP, length 44

10:35:54.667516 ARP, Reply VIRT_IP_3 is-at VIP3_VIRT_MAC (oui IANA), length 46

10:35:54.679669 ARP, Reply VIRT_IP_3 is-at FW1_PHYS_MAC (oui Unknown), length 46

10:35:55.669868 ARP, Request who-has VIRT_IP_3 tell FW2_PHYS_IP, length 44

10:35:55.670034 ARP, Reply VIRT_IP_3 is-at VIP3_VIRT_MAC (oui IANA), length 46

10:35:55.774143 ARP, Reply VIRT_IP_3 is-at FW1_PHYS_MAC (oui Unknown), length 46

  • Here's another arping output for a different virtual IP. This time, the ISP Gateway MAC is responding with the Virtual MAC used for VRRP, which is identical to the local system (primary firewall) response. This particular IP address is pingable from the outside, but can't traceroute beyond the gateway. It hits their gateway and then times out beyond that.

ARPING VIRT_IP_2

60 bytes from FW1_PHYS_MAC (VIRT_IP_2): index=0 time=218.118 usec

60 bytes from ISP_GW_MAC (VIRT_IP_2): index=1 time=9.392 msec

60 bytes from FW1_PHYS_MAC (VIRT_IP_2): index=2 time=185.204 usec

60 bytes from ISP_GW_MAC (VIRT_IP_2): index=3 time=11.139 msec

60 bytes from FW1_PHYS_MAC (VIRT_IP_2): index=4 time=136.903 usec

60 bytes from ISP_GW_MAC (VIRT_IP_2): index=5 time=124.488 msec

10:48:02.686125 ARP, Request who-has VIRT_IP_2 tell FW2_PHYS_IP, length 44

10:48:02.686345 ARP, Reply VIRT_IP_2 is-at VIP2_VIRT_MAC (oui IANA), length 46

10:48:02.695503 ARP, Reply VIRT_IP_2 is-at VIP2_VIRT_MAC (oui IANA), length 46

10:48:03.687297 ARP, Request who-has VIRT_IP_2 tell FW2_PHYS_IP, length 44

10:48:03.687462 ARP, Reply VIRT_IP_2 is-at VIP2_VIRT_MAC (oui IANA), length 46

10:48:03.698416 ARP, Reply VIRT_IP_2 is-at VIP2_VIRT_MAC (oui IANA), length 46

10:48:04.688300 ARP, Request who-has VIRT_IP_2 tell FW2_PHYS_IP, length 44

10:48:04.688423 ARP, Reply VIRT_IP_2 is-at VIP2_VIRT_MAC (oui IANA), length 46

10:48:04.812769 ARP, Reply VIRT_IP_2 is-at VIP2_VIRT_MAC (oui IANA), length 46



CAT 6A patch cables?

Is anyone using cat 6a patch cables campus wide? We've got 6a runs across 75% of our campus but all of our patch cables are cat 6. Is this something I should start to be concerned about or is it still overkill?



Advice needed for POE+ distribution switches

Hey guys & gals,

Need your suggestions for a distribution switches replacement project. About 200-300 employees on site with more and more POE gizmos to power (APs, SIP phones, cameras, blah blah...).

I am particularly interested with the models offering the higher power budgets with robust power supplies. Ideally 48 ports, all of them POE+-capable.

Cisco is okay... but if there is other obvious choices out there that you can vouch for, then good!