Showing posts with label Troubleshooting. Show all posts
Showing posts with label Troubleshooting. Show all posts

Wednesday, August 31, 2011

Win7/Office 2010 Deployment Office 2, Days 1 & 2

"Always plan ahead. It wasn't raining when Noah built The Ark"
Richar C. Cushing

So here we are, delivering Windows 7 and Office 2010 to Nexsen Pruet's second remote office in as many weeks. But first, a little background. First of all, I had not blogged about this project before because I have been in and out of it and my role has been changing. I have been primarily responsible for testing a few applications under my area to ensure functionality in the new environment. However, this week I am the on-site Operations Manager, which could translate into the on-site Project Manager/Do-What-Needs-To-Be-Done Guy (and everyone in the team really does what needs to be done anyway). I am responsible for making sure that the deployments are completed and that issues are addressed and compiled in a central place so that we can go over those before our next deployment. Finally, I'd like to add that this project started very early this year and there have been many phases. We started with brainstorming around January/February, to then go over a Branding phase in March/April. Around April, we began testing our many legal specific applications used by our Attorneys; to then begin creating and fine tuning images. It took until early August and 9-10 versions to get the right image. Developing the outstanding training material took a few months as well and delivering training started in July with Technology Department members first, and then to our pilot group. Last week we deployed the new desktop to our Myrtle Beach, SC office. This week is Hilton Head Island, SC's office turn. The best part of the project? EVERYONE in our department has had an impact, all of the eighteen people that conform our team, including temporary personnel has indeed contributed in a positive way, just as our very supportive and knowledgeable vendors have. But this is about the deployment phase, "The Real Deal". Let's take a look.

A team of six departed Sunday afternoon from Columbia, SC to our Hilton Head Island office. This team includes our Service Desk & Training Manager who oversees the delivery of training as well as floor support; the Operations Manager, who oversees the deployment of the images as well as the many applications and customizations required by each end-user as well as issues logging, one deployment engineer, one hardware and deployment technician and this time, our Director of Technology made the trip to compliment the team and contribute where needed. In addition, twelve others remain at our main office in Columbia holding the fort. Last but not least, I am proud to say that this is The Best Team and Technology Department I have been part of during my 12+ years career.

Day 1. Setup and Deployment

The team got in the office by 6:00 pm and by 8:00 pm the "Technology Operations Center", the Training Room and the temporary Video Conference room were all setup and tested. I have to admit that I thought that we had too many people, but boy was I glad we did! It did paid off. I was amazed of how well and quickly we worked together. This gave the training team the opportunity to go rest at the hotel to prepare for next day's class with staff.

Once everything was setup the deployment team started deployment of images to staff machines and we quickly ran into an issue. We are using Microsoft SCCM to deploy images and we are using our local File & Print servers as local Distribution Points to minimize WAN traffic. Well, the images did not finish deploying to the server. Now, I was out for a week at a Technology Conference and I kind of panicked for a minute and I think that so did our Director, but our Engineer who has done an incredible job on this project even well before we started, was prepared for it and had multiple DVDs with images for each type of machine at hand. I think that we may have lost perhaps a total of 15 minutes or so before we started the deployment. Our engineering and project management teams had prepared a couple of very detailed sheets, one with all the specifics about each machine that we had imaged with information such as specialty applications, printers, drive letters, and anything else that the end-user would need in order to complete their job; and the second sheet contains step-to-step details on how to deploy application and what to check off.

Imaging the machines may have taken about 30 min in average for each. Then we went on to deploy and test most of the applications in each workstation. I tell you, this is very tedious job and what took the longest because of how focused and detail you have to be in order to get everything done properly. We finish around 1:00 am, primarily because we were in such a roll that we just wanted to keep going. After all, we had the whole Monday to complete the deployment ahead of us.

Day 2. The Real Deal

Monday morning I met with the Training team and our Director joined us to go over what we had accomplished the night before and what was left to do before they head to the office to start training the staff. I followed the meeting with a 3 miles run around town, which was very helpful to help me prepare for the day.

The deployment team, including myself, got to the office later that morning while training was being conducted to finish the customization of the machines, which really wasn’t too much after all, the work the night before. Things such as printer drivers, and a couple of specialty apps were the only outstanding issues and they were completed by lunch time. I have to give credit to our Training Manager here; she did a magnificent job in terms of Coaching, Mentoring and Leadership while introducing this incredible amount of change to our users. Want to know why? This was our Trainer’s very first time teaching a Class in our Firm. We did hit a couple of issues but the planning and preparation that went on for months paid up during execution, a perfect example of the 80/20 rule, which if executed correctly does pay off big time.

The day could have not ended in a better way. Our seventh team member, one of our great Service Desk analyst arrived that night to do floor support along with the rest of the team and we head to dinner. After enjoying a great seafood meal at local restaurant, all seven of us took a walk at the beach, which was not only relaxing and needed but it kind of served as a team building experience. IT WAS AWESOME! We looked like seven children having fun!!

Meanwhile, back at The Mothership

Kuddos also to our EVERYONE back at our main office in Columbia. They were so responsive when we needed them and the communication across teams was incredible, another stepping stone for delivering this effort effectively. Our DMS, DBA, AD, Exchange and Network Engineers as well as our main Trainer, Service Desk and Project Management teams were always at hand when issues aroused and were able to quickly address them. In addition, they were also tasked to hold the fort and support day-to-day operations of the firm and they get as much credit as the onsite team does, especially being so shorthanded.

Days 3 & 4 include deployment and training for attorneys as well as support and I will tell you more about it at the end of the day tomorrow, but I can tell you that it was an exciting day for everyone.

Tuesday, June 7, 2011

Troubleshooting WAN Failover with BGP. A good procedure and attention to detail are critical.

A few weeks ago we ran into an issue when our primary Internet Service Provider (ISP) went down and we automatically failed over to our secondary ISP. This had happened before but in a different type of setup when we were in a dual homed Customer Router (CPE), which means that both carriers terminate in the same customer router; today both of our carriers terminate in their own CPE, which are Cisco Routers that terminate in a Cisco 3750 Layer 3 Switch Stack. The main points I want to make with this post are 1) you must have a structured procedure in order to troubleshoot this type of issues, and any other technology problems that you may encounter in your network for that matter. While it took a while to figure out what was wrong with this setup, I believe that it could’ve taken longer and I would have not been able to keep my cool had I not had a procedure to follow. And 2) pay attention to detail, it can save you sometime.

The problem we experience wasn’t really fail over. That did happen as designed. When our primary ISP went down BGP and EIGRP did all their work and the network connection came back up within 2-3 minutes after the failure. I followed our procedure to make sure that fail over happened properly; makes sure we are up (pings and alerting systems), make sure we know which network we are riding (trace routes and alerting system), user experience (make a couple of calls to make sure that systems were accessible), phone system is up (dial some extensions, check PRI registrations at the gateway). Up to this point everything seemed fine. I even called folks and they said everything seemed to be up. However, one person emailed me to report phone issues and at that moment I noticed that our PRIs were not properly registered with our Cisco Unified Communications Manager (CUCM). And here is where the troubleshooting began, not to mention that I am adding this step to the procedure, which is a “living” document.

I ran a few Show commands in both routers and the switch stack and all protocols were up as well as the main routes that we were riding; however, I noticed that my internal routes were not being advertised properly which led me to understand that our voice system was up by means of SRST, the fail over mechanism used by CUCM. I ran a few Show Run commands in all routers and switches in each side. I could tell that Router BGP and Router EIRGP had been told to advertise all the proper networks as shown below.

Remote Switch                                                                                                               Central Location Switch
                    


I then turned my attention to the BGP routers and again, everything looked fine there as well.   

Central Location (Hub) Router                                                                              Remote Router 




I also did run a Show IP Route as well as Show BGP in the routers and Show EIGRP as shown below. I realized that the internal networks from my remote office were not being advertised back to the Hub site but I missed what the main problem was, which I will explain shortly (screenshots below are from current config and they don’t reflect what happened then. Our 10.70.0.0 network was not showing at the moment)






Show BGP told me that I was missing remote LAN network (10.70.0.0 was not showing up)





Some EIGRP Stats



The next step in my troubleshooting procedure was to contact our primary ISP to confirm whether or not they could see the routes that we were sending through our secondary provider, and sure enough, they could not see them. At this point I really felt lost and followed the playbook which says: “Stop wasting time and call Cisco TAC”, so I did. After about 20 minutes of troubleshooting (that after the frustrating 30 minutes on hold), they finally identified what we were missing. CHECK YOU ROUTER ID (RID). The engineer realized that the BGP RID in both routers were the same, hence EIGRP was not able to send the routes properly because it had two BGP routers with the same ID. Once we changed the RID in one of the routers (in our case the primary router) the routes started to propagate accordingly and we were 100% operational while in failover mode. To change the RID go into Router BGP mode and then run the change RID command, where A.B.C.D is the IP address of the interface that you want to designate as the RID.


I am glad that we worked on this until we resolved the issue because this was a very long outage and we were still riding the alternate ISP the next morning. Bottom line: have strong troubleshooting procedures and methods, and revise them as your network changes and evolve, understand what the problem is so that you can tackle it accordingly and pay attention to detail. It can save you time and many frustrations. In addition, I checked other remote sites and noticed the same problem, which I addressed and now all branches are setup properly, which will eliminate further issues.

Wednesday, February 2, 2011

The Impact of Change Management on Technology Operations

For the last 18 months or so I have been researching and advocating for the need of having a better control of change in our environment. I strongly believe that Change Management is not “a thing for developers and project managers”.  It is widely known that many times we cause our own when we upgrade a system, or make a small change to a configuration, that triggers, sometimes not immediately, a sequence of events that can have a negative impact in our network operations and services that we provide. We have implemented a series of tools that have started to have positive impact in our Technology Operations.

Our Tools
The first tool that we introduced was DeviceExpert from ManageEngine, which added serve as our Backup tool for Network Gear and had a Change Management module that soon became popular. Because of this software, I started to understand how Change Management could help us improve troubleshooting, and operations of our network, and I started “evangelizing” about the need for a tool that took it beyond that area and covered as many systems and processes as possible

About six months our Technology Director and I started discussions around how it could help us and we soon added the rest of our Management team and we were all on board. Armed with all the input that we gave him, our Director designed and developed a Change Module that was added to our Resource Management and Systems Catalog portal which we named Techsystems. We use this module to log changes that we make to the systems and processes that we support. Granted, it is a manual process today, but it is already helping us resolve operation issues and it has great potential.

So How Does Change Management Impacts Operations?
In my opinion, one must have feature of a tool like this is the ability to create and schedule reports, and actually look at them. Both of our systems deliver reports that we review individually and then as a team during our weekly meeting. The result? Better picture of what’s happening and possible impact of changes made that were otherwise previously unknown.

I’ll now give you a couple of practical examples of how it helped our team overcome a couple of issues. One day we started receiving blank faxes from our Fax Server, which usually means that there is a communication issue somewhere.  I knew it was a missing setting in either our Cisco Unified Communications Manager or one of our routers that server a gateway to the server, but I couldn’t remember where it was. We spent the next couple of hours going over different settings in the Communications Manager and some of the routers until I finally remember that there was DeviceExpert and shouted it out! We immediately turned to it and ran a Change Management report, which showed the setting that was missing in one of the Routers. Issue resolved. Our problem was not looking at this report earlier because we had never used it. Since then, I have made this my first step to troubleshooting where it fits and I encourage our engineers to do the same. That’s where the value of the tool is. It’s just not a log that you never look at; it is a tool that can help that should be used.

Last week our Financial Systems Administrator found out that our calls had not been properly billed since January 11th. Not good in a Law Firm. He spent a couple of weeks trying to figure out what the problem was because other members of the team had been busy traveling or in projects. This past Monday we finally had time to talk about it and my first questions were to describe the situation and tell me when this started happening. With that information we went straight to Change Management and looked up the log to see what had happened around that time, and sure enough, there were some changes made in the server that host this process. Then, we were able to trace the problem back and revert to original configuration and in about 20 minutes the process was working properly.

If you change it, log it. If you log it,  look at it and share it. What are you doing that is helping you with Change Management?