Friday, December 30, 2011

Reflections on 2011, a year of trial, growth, and questions

It has been a while since I've had time to pick up my bloggers pen.   October is traditionally a hellish month for me, even when work isn't trying, the Cub Scouts have kept me busy all but a single weekend that I recall, and then on Sundays our son plays soccer.   However, things quickly went off the rails this year.   The whole family came down sick over the span of a month, and one week, we were all sick at the same time.  Yikes!  Praise God, we are better now, but that isn't the only change.

My duties on my project have shifted back to more coder oriented tasks, and less focused on testing.  While I enjoy both programming and testing pursuits, I'll admit, I miss the testing aspects of what I was doing before.   It is funny in a way, when I first got approached about a 'Automation testing' position in 2009, I was worried about being pigeonholed as a tester and excluded from some kind of elite club of programmers.  Yet that was a bit a naive thing to worry about in hindsight.  Testing brought back my love of learning in a way I had not felt since college.  It returned to me part of who I always was, but had kept silent in order to make ends meet.  I've learned a lot as I ran that course, and I wouldn't trade the decision for the world.

My current role on project has me pondering though.  I've heard it debated around twitter, about whether you can be both a programmer and a tester.   I know I can do either, but at some point do you not need to decide which to specialize in?  The reality is there are only so many hours in the day for study and growth, and the opportunity cost of each new learning investment, in effect is at a loss for learning something else.  This is a reality that I've now come face to face with in the last two months.   I still have the knowledge from what I learned as a tester, but it has been hard to try and keep up on my learning where testing is concerned, especially when my current responsibilities require me to act in a more code-centric role.

This feeling has left me feeling a bit lost internally because I know I can succeed at anything I choose to focus my efforts upon, it wouldn't matter if it was testing and programming, or some other group of tasks from which I must choose. I have the drive to do what is necessary to succeed.   Still, I find myself at this cross roads because I enjoy doing things that provide value to the teams that i work with, no matter how small or great the achievement may be.  Up till now it hasn't mattered whether I was a tester, programmer, or performing some other service to the project team.  As long as it was value that I added, I was happy, content and felt fulfilled inside,  Yet I find myself feeling as though I am stuck at a fork in the road.

I feel as if I have paused at a great fork where two rivers meet.  One is a possible focused career on testing, the other, a continued focus on programming, its methodology, and potential as a generalist programmer.  Either fork in the river looks potentially enjoyable from a learning stand point, with its opportunity to pause to fish, relax, or just skip a rock to the other side.  Like most rivers though, I realize that I can only paddle up one stream at a time.  Although the left fork might be easier, the right might be the more fulfilling, or the converse could be true.

For nine years, I have worked professionally to develop, test, and support various software efforts.  I have learned something from every experience that I have been fortunate to endure.  I wouldn't trade those experiences away as they define a bit of who I am personally and professionally.  As I enter my tenth year of service in software development, I find myself looking back over the peaks and valleys behind, and ahead up the forks in the river, yet seeing behind the first bend of either is impossible.  So I am presently anchored, where I am at this fork, pausing to consider and reflect upon what my dreams are for the next ten years.  Where do I want to be?  What roads will I need to travel to get there?  These are questions that I have no answer for currently.   So given that the new year is around the corner, I can see myself at least initially focusing greatly upon what exactly it is that I most want to do, and the realities of that choice which may require not just myself, but my whole family to adapt as well.

It may take some time for me to come to some answers, and the pot has clearly become foggy and hard to see how its contents will turn out when I finally reach that conclusion, but I want to consider things more closely, set a plan and then rush after it to attain it.  Perhaps it is the nature of how I 'fell into' my current assignment that is at the heart of this muddled mind of mine.  At least I know it is something I can do for now, while I sort through my feelings and make what could possibly be the biggest personal, and professional decision I have made in my life thus far.

But enough about me, as you read this and other blogs, I imagine you may be reflecting on recent events, just as I have been.  Where do you stand?  What's your dream?  How will you decide what to focus upon this year, and as a result, what areas that get left behind will you perhaps miss when we reach this point a year from now?

Saturday, October 1, 2011

Diary of a Soccer Coach: Week 4

You've been there before a well worn meeting room with your team gathered around a table going over a list of action items related to the project you've been working.   Sometimes they are new requirements, perhaps they are refinements of existing functionality, or tweaks of the deployment procedures taking into account lessons learned in that first deployment of the software.   The first practice after that first game,  is very much like this.   Often there isn't enough time to discuss defensive or offensive tactics with the players before the first game.

It is sometimes a matter of perspective, a coach after the first game of the season, or a manager or lead on a team reviewing the steps they took on that first ever critical deployment, the first game, the first actions of substance as far as the customer might see.  So in hindsight, that first practice, that first meeting, is often a discussion of the aftermath.  For our kindergartners, we discuss the issues I noticed during that first game.  There are almost always a few areas to correct, and they aren't always the same from Game 1 in one year versus any of the others.

Typically a reminder of the rules is necessary.  A reminder about which goal we are attacking, which one we defend, a reminder not to use hands except for the Throw-ins, and an encouragement to stop play when the whistle blows and quickly bring the ball to the referee when there is a stoppage of play for going out of bounds.  While errors can occur in any game, I try to point out the mistake, and correct the behavior without singling out any particular team member.  The point after all is not that someone erred, but that we play the game correctly to reduce stoppages of play.

With the instructional league sometimes this is difficult.  Some players have an over arching desire for the ball, and they may indulge this by diving at the ball.  This is a behavior we try to discourage.  For one thing, falling to the ground is as bad as waiting flat footed for the ball.  They aren't upright able to move with the ball, and if they are down where the ball is, there's a higher possibility of injury as other players go for the ball around them.   Sometimes they fall down, and stay on the ground, and again even if the ball isn't near them, this isn't behavior we want to encourage.  Sometimes its a sign that the kid is tired, but they rarely will get into shape if they sit down when they are supposed to be in the game, and the ball could come at any time too.

Ever been on a team where similar behavior happens?  Where they get tunnel vision, seeing only the one task before them at the expense of what is going on around them on the field?   How can we avoid this behavior?  What's worse is what if a team mate ends up blocked or stuck in some area and doesn't realize it?    As professionals we can try to encourage, give a second set of eyes to these issues, but in the end its really up to the individual to get themselves back on track.  We can encourage that team mate to come back on board, but honestly, if that person cannot take the initiative, there may be little we can do to really fix an issue that is internal to them.

Just like with my soccer players, I take them back to basics, breaking down the basics of the pass, of corner kicks and goal kicks, of throw-ins and kick-offs.  Only so much time can be spent on correcting the past, as new challenges and new games await.  So before the second game we spend a little time talking about defense.  Of reminding the kids that in our league, there are no goalies and therefor noone should go into the goal arc even to go after the ball, but more importantly we teach them what to do to protect their own goal.

First involves positioning, if a player is following an opponent bringing the ball up, we show them how they can move their feet with out crossing them, using the balls of their feet to have better response time in their jockeying back and forth.  We show them how to encourage the ball handler to dribble a particular direction, to funnel them away from a straight shot on goal, or to where we hope additional team mates can cut off their lane of advance.  We also try to show them that having everyone covering one person leaves open lanes of passing to the opposition, it leaves area of the field uncovered, and opens up easy attacks on their teams goal.

In software development, testers play a part of defense, not from bugs scoring on them, but from preventing threats to the value of the product.  If the goal as a team is to release a product with value that's usable by the client, then anything that allows the product to be misused, leaves features less than fully implemented, or just plain not covered is a threat that we as testers try to find.   The one difference here is that unlike in soccer where we can see the ball coming many times before it arrives near our zone of defense, in testing we don't have the ability to look at the software and say a bug is coming from here or there.  We have to instead visualize it with our mind.

How can we visualize where bugs might be?  One way is to be involved early in the process, be in with the conversations with the customer or client and helping to determine how the software may be used.  We also must consider the negative, the view of what invalid data, or improper operations might do to the software.  What if a file consumed is missing settings, does the software resort to a default and store that in the configuration for next time?   If you start typing before the software can fully load, will it cause an unexpected behavior?   We can brain storm a horde of test ideas to try to cover the entire areas of the application, but the reality is just like in soccer, we are just one tester, we can only cover so much ground in eight hours of work time.

What about opposition tendencies? It may be possible in soccer to see that certain players tend to favor an attack on goal from the right or left side.  Some players may prefer passing the ball forwards, or looping back rather than continuing forward at a bad angle.  As testers, we can evaluate the software for tendencies, are there certain areas that seem more bug prone, are there areas that are more critical, or more likely to be highly used and thus could cause more risk?  Is there a particular feature set which sets your software apart from another, then that is an area I'd be sure to test.

Then a foul may be called.  Maybe one player pushed or tripped another, maybe it was a hand ball.  Maybe there's an area of your software that is of particular risk to the customer.  They need that feature to work, quickly, to solve a time critical problem.  Whatever the case may be, we try as hard as we can to find every single bug there may be, but the reality is we can't cover the whole of a software that's anything but trivial.  The nature of software and the myriad of systems it may be installed upon create such a large volume of possibilities that we cannot test it all, so we use techniques to break the software down into areas that we can cover.  We find ways to distill problems to a range of possible outcomes, and we try to think of new ways to test old functionality, because you just never know when a new feature may impact an old one.

Tuesday, September 27, 2011

If its not random, how to decipher the pattern?

Earlier this week, I wrote about the software fault in the Staunton, Virginia, teacher payroll system.  I talked at length about the concept of 'random', and the importance of distinguishing between something that which is truly random, from something that is better described as unexpected, unpredictable, or just 'having no discernible pattern to me as far as my sense go'.  Using precise language when describing defects in software benefits everyone on the team, including the customer.

Unfortunately, the fault in Staunton, Virgnia's payroll system wasn't found by a tester, instead it was discovered by someone researching the finances of the county's school system.   Now we may not know exactly how this fault was first brought to the attention of this school district.  That doesn't preclude speculation how an investigation of a similar fault on a hypothetically similar system could be conducted.

So imagine a hypothetical payroll system for an organization with multiple locations, accounting for user entities of diverse pay grades and positions similar to the school systems.  The more layers you add to the structure of the system, the greater its complexity.   Now let's suppose the vendor of this software received word about an apparent bug.  This bug affects certain persons within the system who would receive an unexpected, and here to fore unnoticed pay increase.  So if you as a software tester for this software vendor, receive this notice where would you start?

Reporting and analyzing a defect that a tester stumbled upon through his or her own investigation of the software is one thing.  Trying to track down a flaw someone else found and reports is quite another. If we follow the example of the of the system we discussed earlier we can imagine the reports taking the form of output, potentially pay stubs, ledger logs, bank statements, etc.  In short, we possess a log or evidence that the problem occurred, but this evidence may be far enough from the system itself to not be able to produce the same conditions without a bit more digging.

So how can we reproduce these conditions and figure out where the real defect resides?  More information is required, and like a software Sherlock Holmes we must examine the evidence, and piece together the story of what happened.  In the case of the pay roll system it is likely important to know how many individuals were impacted.  Might a search for more information related to the individuals effected, reveal each user to be part of particular entities or organizations with in the system?  Did they work at particular locations, or have their data maintained at a particular data center?  An exhaustive analysis of whatever data can be culled from the system could help establish a definitive relationship between the affected users.

From the headlines, it sounds the School district performed an analysis just like this.  The result seemed to have something to do with individuals who went to a particular school.   Now the age of the defect in the system may not be clear.  If  a lot of time has gone by, it may be possible that the connection is more subtle, and won't track to any particular organization, or be so obvious; however, in this case we strike pay dirt.   One piece of the puzzle is in place.

Given that all those results might track to a particular organization within the software, this may lead to our first hunch.  Were all the people assigned to this organization, also receiving the same bug?  This might be where the first bump in the investigation may be encountered.   Maybe they aren't all affected.  That idea may lead to a belief that our initial hunch was wrong, but it could be that there's a reason why they turned out to be the exception.  

It's at this point in the defect analysis where a history of debugging similar enterprise applications could prove beneficial.  From reviewing some of the articles around the defect, a number of ideas come to mind, all of them based on similar behavior I've encountered in other projects I have worked.  If these employees all worked at a facility that was shuttered, what happens to their accounts when the facility is shut down?  Are they transferred to a new facility?  Are they suspended out right?  Are they removed from the system?

I recall once with a customer relation management system that we encountered a bug when a user account was removed from the system.  All the records linked to it, would cascade and delete, or disappear and not show up in the system when searched.  Could a data integrity issue regarding the integrity of the data for these closed locations be responsible for this behavior?

Another possibility that occurs to me, is that a system that freezes pay for all employees may apply to a group of employees by group.  Might a group that these employees belonged to be used to freeze all of their pay for some time period?  Might failing to belong to a group due to the original group being inactivated cause the issue of applying this freeze to miss these accounts?

It may be difficult to see the cause from just reading the few reports you receive from the user, but a simple logical, and step by step examination of the system could help reveal how the issue happened, and if it was a case of the system being used in a manner that was unplanned by the software vendor, it may indicate a fault in the business rules, or lack of training for the users of the system.  Whatever the case, the team is now on its way to finding where this issue occurred.  Where would you test next?

Wednesday, September 21, 2011

Pay Freeze, slightly melted, a random bug? Maybe.

Every now and then I read about a problem in a software system that makes the news.   I look at the article and read what is described as the problem, and I often wonder, how this supposed flaw got into the system.  In my experience it can be easy to fault the software for an error.  There have certainly been enough cases of odd failures for the general public to believe them, but is it really the software?

This week I heard about the story from Staunton, Virginia.  Apparently the school board had frozen pay for all of its employees for some period of time, and as the article stated, the glitch went uncaught by a number of employees who spot checked this up until a news station requested records for salaries under the freedom of information act.  This is when the discrepancy was apparently noticed.  Now this glitch appears like something of a scandal.  The political black eye alone could be enough to make anyone nervous about the 'quality' of the application in question.

What concerns me though is that this glitch is being described as completely random.  First, do we really understand what it means when something is truly random?  According to Dictionary.com, random has four customary definitions.  The first means 'proceeding, made, or occurring without definite aim, reason, or pattern.  The second is its use in statistics, specifically the concept of a process of selection whereby each item of a set has an equal likelihood of being selected. The third definition applies to physical trades, where a part, parcel, or piece of land or item may appear non uniformly shaped.  The last one is an informal use implying that it was an occurrence that was completely unexpected, or unpredictable.


Let's take a moment and consider the story for a moment.  The first definition implies that there is no rhyme or reason, no discernible pattern to something which may make it random.  Is that the case here?  Reading further I notice the following:
" the pay increase malfunction was random and included three teachers at Bessie Weller Elementary School, four at McSwain Elementary, four at Ware Elementary and a speech teacher and a secondary special education teacher."
Several of these teachers had one thing in common, they attended one of three elementary schools nearby.  Wait does that mean what I think it means, could this be the beginning of an actual pattern emerging, enough to discount the perceived randomness?  It could be, but as testers in this situation, our job is to determine the nature of the fault, not just give our 'best guesses'.  We know a fault happened, therefore we must find a way to duplicate it.  If we continued on this analysis, we'd likely have a couple of test ideas to begin testing, we'd look at the data for all of the affected persons and see just what is it that happened.  Is the over payment of salaries here the problem, or is it a side effect of some other hidden flaw that just became visible due to some quality of the instance that we are examining?   Fortunately, I did a bit more research and found another article on this on MSNBC's site.  Now I will note that MSNBC's article is dated the sixteenth of September, and the other article earlier on the Second day of September, however, as I read I find another nugget that seems to confirm my suspicion.

"All the affected teachers had previously worked at Dixon Elementary School and were reassigned to other schools after Dixon closed two years ago."

So it appears that this bug affected teachers that had all been assigned to a school, that closed two years ago. (No doubt around the time of the glitch actually occurring.)  Would you call this random?  No I see a pattern, so it doesn't hold on the first definition.  The Second definition doesn't hold up to the story at this point either, as given a sampling so large, would you really expect to find just a handful of salaries that are wrong?  I don't buy that either.  The third definition doesn't apply in this context, which leaves us with the remaining informal definition: simply that it was odd or unpredictable 

This fact I do not doubt, no one predicted this to happen.  Now I'm not writing this to criticize the vendor or the county in question where this happened.  That's not the point of this article.  Instead, my hope is to make you think.  As a tester, developer, user, consumer of computing appliances, how often do we encounter behavior that surprises us?  How often do we not only get surprised but feel the event to be unpredictable, with no reason it should be happening?

I imagine this happens more than we might like to admit.  How many times do we sit at our computers, doing something normal.  We're checking our email in our client of choice, we have had no problems with our service and expect to get a no messages found if the service has none waiting for us.  We hit the send/receive button, and wait gleefully hoping to find/not find email.  Then we get a message that it was unable to connect to the server.   That catches us by surprise, maybe we think its an aberration, so we click the button again.

That second click does what?  It allows us to check to see if it was a hiccup, a momentary failure, or perhaps a sign of a long term issue.  I've had this happen from time to time on web pages I may visit frequently.  A forum for a football team may load very fast during the week, but on game day as people are checking up on their team, it slows to a crawl, and a dependency like a style sheet, or images fails to download due to the sudden hit to bandwidth serving the multitude of requests at one time.  It might even take minutes before you get that white page with some structure, and no formatting.  Do we immediately think, wow that's random, this forum is really bugged?  But I know from experience, this isn't a fault of the software itself, at least as far as I can tell, but instead it is a function of a high load on a system that may not be able to keep up with a sudden increase in demand.

As testers, simply finding and reporting bugs is wholly insufficient to communicate to the developer the nature and scope of the fault we've encountered.   In the case of the forum software, a subsequent refresh might fix the page, and it may load fine for several hours thereafter, unable to have the issue reproduced.  Whatever the issue is, we must dig, and see if we can prune down the steps that we followed.  We can try to see if the bug happens if we hit another location, try a different path through the software, or perhaps try a different role or persona.  The point here is it is our job as testers to imagine how this bug could have occurred.  What would your tester instincts tell you to go to prove and find this error so it could be fixed?  Do you have the answer?

Hold that thought, because I am going to revisit this question later in the week.  For now, just remember that just because we can't see the pattern for a bug, doesn't mean there isn't one, and as testers in particular, our use of language should be careful so as to not mislead the public, our developers, our clients, or our managers.

Tuesday, September 20, 2011

Diary of a Soccer Coach: Week 3 and First Game!

I'm a bit behind on blogging due to duties last week, but I'll catch up by throwing the third week of practice alongside the first game.   In our league the third week of practice heralds two things, first the last practice before our first game, and the arrival of our team rosters and uniforms.  On this particular day, the 'head coach' of the league, who has been helping out with our Kindergartners had to distribute the uniforms to all the different divisions.  This left me alone to do a lot of work with the kids on my own.

This worked out fine, and as I always try to keep the kids moving it worked out great.  The Third practice is where we honed in on basic shooting skills.  For most introductory soccer players, the more advanced steps are not always easy to pass on.  At this practice I focused in on keeping their eye on the ball and following through as they shot.   Also, as with most practices I got and kept them moving as much as possible.

I started by having them stand next to the ball and shoot it stationary.  After each player had tried this a few times, I had them try shooting the ball by first running up on the ball and then kicking it into the goal.  Afterwards, I made it more difficult by having the players dribble the ball and then kick it into the goal.

On many development projects I've seen a similar step by step building up to completion.  A feature might start out very simplistic, or it may seem that way so we start by taking our first shot at it, just as my players might in their third practice.   Sometimes we may not understand some of the nuance to a requirement.  It may appear simple, just like striking that ball, but there are intricacies and un revealed flavoring that needs added for the code to really pull off what is intended.  So as a team maybe you work up to these features, adding a little more speed, a bit more control, and higher accuracy in its calculations.

I've found similar patterns in testing.  The first time through, you may just be poking around in an exploration of the application under test.   You may not have a full grasp of the features, how to activate or use them,  or the intent, but you build a bit of confidence and then take another test run at the software.  Then you might discover that this type of software is documented to have a particular susceptibility to one kind of fault, and begin tailoring your exploratory testing to hit those weaknesses.


The first game of a soccer season is always exciting.  Its the first time the kids are in their new uniforms, and you just never know how much the kids have absorbed from the limited practices you've had thus far.  Each year is a little bit different. One year, one team may have a very good grasp of the game and create a lot of goals in that first game.  Others might find it difficult to juggle defending the approaching ball, redirecting it to the goal they are attacking, or they may even get a little winded as they aren't used to moving so much at one time.

The first year I coached, an older coach told me, "You'll see the most improvement between the Second and Third games." I wasn't so sure how to take that, but later I realized what he meant.  Suffice it to say, many kids may not listen early in the practice.  Until they see how they can apply it in a game situation, they just may not realize the advice you are giving them.  I've seen this happen on development teams too.  A tester might make a suggestion about how to improve a process or function within the application, and might be ignored, because its simply not their job, or because the developer is too much ' in the zone' to stop and see what is being said.  There might even be, as is common in our first game of soccer, a lot of stops and starts as you build to a sustainable pace for development.

Bottom line though, remember it's just the first game.  A lot can change over the course of time on a project.  Change is inevitable in many projects, and how we handle and respond to it sets a strong light on our teams and how we cope with that change.