Tuesday, October 21, 2008

Bursty

That's always been my M.O., I guess, but last week was pretty busy:

Flyway

I went down to the Flyway Film Festival in Pepin, Wisconsin over the weekend to screen "Cottonwood", a short film I directed over the summer as part of the 48 Hour Film Project.


I gave a brief Q&A about the making of the film, then did some schmoozing at the filmmaker's reception. I had a great chat with the director of "
Redneck Zombies"--the guy was hilarious.

"Cottonwood" is five minutes long--six with credits--and was written, shot, edited and scored in exactly forty-eight hours. And although "Cottonwood"
looks like a film that was made in an absolute, mad panic (it was), we had a blast doing it.

Minnesota MySQL+PHP Meetup

Last Wednesday, I had the great pleasure of attending the local MySQL+PHP meetup.

The attendance was excellent, the attendees were interesting, and I had the opportunity to talk about Falcon, give a rundown of the Riga meeting, and dispel (I think) some of the rumors making the rounds in the MySQL community.

There were many questions about the Sun acquisition and its effect on MySQL. I was happy to emphasize how positive the acquisition has been, most recently evidenced by the addition of several developers to Falcon and a commitment to avail us of the PAE performance team's expertise.

I also met three local MySQL customers who are interested in giving Falcon a test run with their web applications. When we're closer to beta, I'll arrange to visit them and help with the configuration.

Chill/Thaw Madness

I spent the rest of the week and weekend completely immersed in debugging several chill/thaw and concurrent online alter bugs emanating from the System QA stress tests. Falcon must be stable before we can hand it over to the PAE team for testing--that has been my number 1 priority since Riga.

I have a pile of fixes to push--some of which I will do tonight--but after wrangling with several nasty thaw bugs, I've come to realize that any mechanism that complicated should probably be scrapped (more on this as it develops.)

Ok, now I'm getting agitated. Back to it.

Some Perspective on Recent Events, Part II

"Why Falcon Doesn't Work Anymore"

Recap: After the MySQL Dev meeting in Riga, Philip "Troublemaker" Stoev, reported a serious problem with Falcon recovery. Ann suspected that the problem not a true recovery bug, which Jim confirmed.


The problem wasn't that we found a bug. The problem was that we didn't.

Run-up to Riga


The Falcon team was really rockin' in the last half of August and early September. With the Dev meeting looming, all changes for the 6.0.7 alpha had to be pushed by month's end, otherwise we'd lose weeks of QA testing.


Lots of stuff was dumped into the codebase during that short time. The Falcon weekly report for August 22 records pushes to the Transaction Manager, Memory Manager, Page Cache and Online Alter, just to name a few.

The build team finally cloned mysql-6.0-falcon on September 4th, but the heavy volume of source changes continued, peaking on the 10th, one week before the conference:

22:57:47 <wlad> Today is the absolute world-record in number of pushes/hour in the Falcon tree
22:57:47 <wlad> I think we already had 8 and seeing Chris' determination, there are more to come:)

22:58:56 <cpowers> Doing my best...iterative process of refinement, you know...
22:59:19 <cpowers> Got to keep them pushbuild servers happy
22:59:27 * wlad applauds:)


Pushbuild is the MySQL automated QA system. Every code push into the source tree triggers a wave of regression tests across twelve uniquely configured servers. Test results are displayed within in a simple matrix, where each cell represents the status of a single code change on a single server. If the cell is green, all tests passed on that server. Red indicates a test failure, build break or even a compiler warning.

A high number concurrent pushes will cause Pushbuild tests to stack up, creating a multi-hour delay before results are available. Such was the case during the Great Week of Pushes as the Falcon team checked in change after change. The frantic pace also resulted in several preventable regressions that kicked the tree back into the red, necessitating repeat check-ins and further aggravating the Pushbuild backlog.

However, by the end of the week, Pushbuild returned to its refreshing, reassuringly green self, and we prepared for our journey to Latvia.

Worklog #4479

During the pre-Dev meeting melee, one item made it into Falcon tree with relatively little fanfare: a two-part cache optimization provided by our performance expert, Kelly Long. From the associated worklog descriptions:

Falcon page cache has a findBuffer() routine that looks at the BDB list starting at the oldest and moving toward the youngest. This code is single threaded and can include time consuming disk IO. This change will increase the concurrency by executing most of the algorithm without locks.

Falcon page cache currently uses a single lock on the entire hash table. This change will create a lock per hash bucket allowing for higher concurrency.

The general idea was to push locking further down into the cache by replacing a single exclusive lock at the top with separate locks for each slot of the hash table.


Jim rejected the notion when it was first proposed back in May, pointing out that the algorithm would effectively triple the number of exclusive locks, even for cache misses.

Kelly was convinced, however, that the change would allow for better parallelism, and that a loss in performance at lower concurrency would be more than compensated by a performance gain at higher concurrency. His own benchmarks appeared to support his contention.


The issue was revisited--contentiously--at the July meeting in Boston, but with no consensus. As the Dev meeting grew near, Kevin reconsidered the matter and eventually granted approval for the change, so Kelly's cache optimization finally made its way into the Falcon tree--along with twenty-five other changes in the same week.


Hard Landing


The Falcon team returned from Riga at the end of September, tanned, rested and ready for the drive towards Falcon beta. No sooner did we sit down at our dusty keyboards than did Philip alert us to new problems with Falcon recovery.


Recovery is a normal part of the engine initialization process that reestablishes the previous in-memory state and applies pending updates following an abrupt or abnormal shutdown. Recovery failures are serious and notoriously difficult to debug, so Philip's discovery of new failures stirred things up considerably.


Fortunately, Ann had a hunch as to what the problem might be. After some digging, Jim pinpointed the failure as a bug in the oft-debated cache optimization:

"The problem is simple and straight forward. Cache.cpp is totally broken. More specifically, the I/O thread (and architecture) is totally broken."

"When a checkpoint was required, a dirty page bitmap was produced. The bitmap, however, didn't contain the table space id, so ioThread had to loop through the collisions in the hash table to find the dirty buffer(s)."


"This only handles the first Bdb in a collision chain. In almost all circumstances, this will be the falcon_user tablespace. Consequently, dirty pages in falcon_master are never written."


"I'm sorry it has taken this long to discover that Philip's problems were page cache bugs masquerading as recovery bugs."


"I think we may need to reconsider some of our procedures."
The fix appeared to be simple enough--ten minutes by Kelly's estimation--but rather than spin more cycles on optimization, Kevin chose a reliable but somewhat heavy-handed course of action:
"In light of this analysis, I think it is prudent to back out Kelly's cache changes. We need to focus on delivering the most stable Falcon engine that we can, most immediately, to the [performance] team, who are our #1 customer right now."
The performance team was the Sun PAE group lead by Allan Packer. During the Dev meeting, we agreed to get a stable version of Falcon to them ASAP so they could begin their own in-depth performance analysis.

But the issue at hand was much larger than a code regression--we'd certainly had them before--the issue was that a serious bug escaped detection.

Vlad got right to the heart of the matter:

"I do not see any way to prevent 'totally broken' in the future. How should we should change our processes to detect an error before the push? It passed all the system tests."
Indeed, the problem wasn't with the Falcon code, it was with the Falcon process.

(The third and final installment the story continues tomorrow, "Perturbations in the Field")

Tuesday, October 07, 2008

Sort of Out of Sorts

I promised Giuseppe that I'd post Part II of "Some Perspective on Recent Events" today, but I got hit w/ some kind of bug.

Last month, authorities at Schiphol duly confiscated my Black Baslam, aka "The Latvian Hammer", so I've had to resort to bland, marginally effective over-the-counter remedies.

Let's see how tomorrow goes.

Monday, October 06, 2008

Prolificity

I woke up today at 4:45 am, which is a completely worthless time to be awake. I lay there for twenty minutes tracking post-REM phosphenes and thinking random thoughts. Here's one of them.

The Planet MySQL aggregator ranks blogs according to 'activity', or number of posts in the last month. I typically care little for such things because unnecessary attention attracts unwanted expectation. However, were one to care about such things, this method of ranking favors those who tend to post like this:

Sunday, October 5
I broke my pencil but that's ok because I have another pencil.

Monday, October 6
I like Drizzle.

Tuesday, October 7
I like, drizzle.

Wednesday, October 8
I believe I lost my shoes, Clyde. I think the dog got 'em.

Free idea: Rank blogs using a Prolificity factor that incorporates two subfactors, Currency and Abundance.

Currency is the number of posts in the last month, and is a fair measure of how current the blog is.

Abundance is the moving weighted average of word count per post, which is a measure of the substance of the average blog post.

A moving weighted average will favor more recent posts, so that 100,000-word post of your PhD thesis back in 2002--though a fine contribution to All Things--would have little relevance today.


Expressed formulaicly,
Prolificity = Currency*Abundance
Let's say Currency wins in the event of a tie, so maybe this:
Prolificity = Currency*Abundance + Currency
For example,
Blog Theta: 10 posts, 2000 words, 200 words/post MWA last 10 posts

Currency = 10, Abundance = 200
Prolificity = 10*200 + 10 = 2010

Blog Zeta: 2 posts, 2000 words, 1000 words/post MWA last 10 posts

Currency = 2, Abundance = 1000
Prolificity = 2*1000 + 2 = 2002
Blog Theta ranks because it's more current, but Blog Zeta isn't entirely out of the game. Perhaps the formula could be tweaked, but I'll leave that as an exercise for the reader.

Sunday, October 05, 2008

Some Perspective on Recent Events

[This has nothing to do with the financial crisis or the McCain/FeyPalin ticket. Sorry.]

Last Wednesday began innocently enough. Philip's email to the Falcon team, "Regarding Falcon Recovery", lamented the lack of progress in fixing recovery bugs. He detailed specific failures that he was seeing, many of them new, and pointed out that that the number of recovery bugs was increasing. He closed with
"All of this means that the recovery problems must be tackled immediately and head-on. A database without functioning recovery can be at most alpha-quality software."
He was, of course, absolutely right. The lack of reliable database re
covery is like flying with non-locking landing gear, so we took a hard look at the problem.

Then all hell broke loo
se.

Falcon Recovery

Falcon recovery is the equivalent of waking up after a night of partying and trying to figure out where you are, what you did and who you did it with (I have never experienced such a thing. Ever.)

When Falcon starts, it determines whether it was shutdown cleanly. If not, it invokes recovery to apply the remaining active parts of the serial log, beginning with the last serial log entry, then backing up to its recovery start point.

This multi-phase procedure must work without fail, every time, even following an aborted recovery attempt (this is called double recovery.)

The recovery process applies changes in the serial log to pag
e images in the page cache, until they are all done, then flushes the cache by forcing a checkpoint. Falcon does this in three passes through the serial log:

Phase I: Take Inventory, Establish State

- Determine transaction states
- Determine object states (track state transitions, record final state)

- Determine the last checkpointed record (prior objects guaranteed on disk)

Phase II: Physical Allocation
- Allocate and release required pages and sections
- Track object "incarnation" and state--update active objec
ts with last incarnation

Phase III: Logical Application

- Apply data and index changes (avoid reallocating pages in use)

(For details, have a look at SerialLog::recover() in SerialLog.cpp.)


Recovery failures are notoriously difficult to debug. Sometimes you have to deter
mine not only what did happen, but also what didn't happen, such as a missing page update. And quite apart from tracking various types of recovery objects, recovery must also track the state or "incarnation" of each recovery object. For instance, if a record was updated five times, we only care about the most recent version--the fifth incarnation

(Free idea: "The Fifth Incarnation", an epic tale of lives lived, lives lost, loves lived and lost loves. At quality book stores near you.)

"Expected 7, got 2"

Philip modified the SystemQA stress tests to force a recovery and found that the engine crashed every time. Up to this point, Falcon recovery was regarded as generally reliable, though not flawless, but now we had a reproducible failure that always manifested the same way:
Fatal assertion in Cache::fetchPage(), "Error: page 103/0 wrong page type, expected 7 got 2".
Cache::fetchPage() lies within the very nucleus of the Falcon database engine. It is the cosmological core through which every essential database action must flow.


Getting a "wrong page type" assertion in Cache::fetchPage() is like asking your HAL 9000 to compute "4 + 4 and getting "green" for an answer. This is not the type of error you ever want to encounter in a database, or in Jupiter orbit for that matter, because it puts everything into question. Nothing can be trusted--data pages, index pages, orbital trajectories--nothing.

Given that the cache error appeared to occur only during recovery, and given the c
omplexity of the recovery process, it was only natural to assume that we had somehow failed to properly record a page update within the serial log. (We later discovered that the error had indeed occurred elsewhere, but only occasionally. Hakan saw it during DBT2 testing, and I had also seen it testing online alter.)

"This used to work. Always"

Ninety minutes after Philip's email, Ann responded:
"This appears to be a recently introduced problem that (probably) has nothing to do with recovery. It appears that the PageInventoryPage is not being written often enough. I don't know why. As a result, pages that are actually in use are recorded as being free. This isn't a problem for a running Falcon because the page is up-to-date in the cache. After a crash, the out-of-date version is read from disk and pages are reallocated."
Falcon uses PageInventoryPages (PIPs) to track used and available pages, so an out-of-date PIP implies a very serious corruption.

Jim provided his assessment shortly thereafter:
"It doesn't work at all. There has been a huge regression. The simplest of all possible test -- start with a clean database, create a single table, kill the server, and restart -- fails first time, every time."

"This used to work. Always."

"We don't have a recovery bug, we have a non-functional code base."
And so it began.

(Part 2 of the story continues tomorrow, "Why Falcon Doesn't Work Anymore")

The Weekly Falcon Index

Planned Falcon Blog posts: 4
Actual posts: 1

Team emails debating the cost of changing core components: 44
Post-debate emails suggesting a major change to a core component: 1
Days advance notice to change my LDAP password: 9
Nag emails to change my LDAP password: 9
Days of delay before changing my LDAP password: 9
Hours unable to access email after changing my LDAP password: 6
Calls to support after changing my LDAP password: 1
Falcon IRC non-system messages: 2022
IRC champ: Hakan (23%)
Runner-up: Vlad (21%)
Messages in German: 35
References to "Ramadan": 10
Best quote: "I try never to argue against clarity"
Source of quote: Ann
Intrusions by senior staff: 1
Number of times this week I typed a password into IRC: 1
Number of times this year: 3

Monday, September 29, 2008

Falcon in London

[I'm still playing catch-up after a five-month blogging hiatus]

By May of 2008, the Falcon team's technical center of gravity lay squarely within the European continent, so it was only fair and sensible that we gather where the combined travel effort would be at a minimum.

This was to be the first assemblage of the entire team, including new team members. We were also promised (and later cruelly denied) a rare sighting of Hakan K. who was unable to attend the all-hands staff meeting in January or the UC in April, as well as appearances by MC Brown and James Day.

MC did an outstanding job as a one-man advance team, deftly applying his hitherto unknown logistical genius to the task of arranging for accommodations at a reasonably-priced, centrally-located, well-appointed and distinctly British hotel..."Holiday Inn" I think was the name.

Big News

We started the week in London on an interesting note: Jim announced that he was stepping down from the Falcon project in order to pursue his cloud database idea. This wasn't a complete shock, because Jim had been itching to explore his cloud database concept for some time, but it was still big news.

The timing of the announcement was pretty good, because Falcon was nearly feature complete and the team had gained enough expertise to carry on. We were also relieved to find that Ann Harrison (Jim's wife) intended to remain with Falcon, because she has a holographic knowledge of much of the architecture and would provide a reliable back channel of communication to Jim--reliable because they share the same office and that she can, and does, holler questions to him over the partition.

We also had a full-time project manager and three new engineers, so Falcon was clearly in capable hands. The only remaining issue was whether Jim could arrange a deal with Sun to continue assisting the project (which has since been worked out, so not much has changed on a day-to-day basis.)

Falcon Immersion

A key item on the agenda was Falcon training for the new team members. Jim provided an excellent, in-depth review of the salient aspects of the Falcon design, reinforced by line-by-line walkthroughs of certain core operations.

I took notes which will ultimately make their way into Falcon design documents and even the code itself (the stringent use of source code comments within Falcon warrants further discussion, as does the overall coding style. I will address this in a future post.)

It was a long week of long days, and we covered a considerable amount of material (I will also cover some of these in future posts):

  • Record insert process, from StorageInterface to data page
  • Serial Log record formats
  • Page Cache design and areas for optimization
  • Internal SQL engine
  • Low memory operations and record cache maintenance
  • Falcon on-disk structure
  • Falcon indexes, design and operation
  • Page types: Data, Index, Inventory, Overflow, Blobs
  • Engine initialization sequence
  • Falcon handler and the MySQL Storage Interface API
  • Overview of the Falcon Storage classes (StorageConnection, StorageTable, etc.)
  • Synchronization objects
  • Replication: Statement-based vs Row-based
Statement-based Replication

So let's explore that last item a bit.

You know how it goes: Someone asks an innocent question in a room full of people, and before you know it, an hour of unintended discussion ensues. Such was the case with Falcon and replication.

The question was: Why doesn't Falcon support statement-based replication? The complete answer is complicated and worthy of a separate post, so I will only summarize here.

Falcon supports row-based replication (RBR), where every row event on the master lands in the binlog and is subsequently pulled in by the slaves. One advantage of RBR is also it's greatest disadvantage, that is, everything is logged, which is fine except for unqualified statements like "DELETE FROM T1", or for operations that are later rolled back.

On the other hand, statement-based replication does not handle non-deterministic or concurrent operations very well, and must therefore rely upon the assumption that two statements won't have any interaction. Ann puts this very succinctly:

"Falcon does not and will not support statement based logging because Falcon transactions are not serializable, and in the absence of serializable transactions, statement based logging does not produce consistent results."

As I understand, InnoDB gets around the problem by locking record touched by subqueries, obviously at the expense of concurrency. For Falcon to support SBR, the same would also hold true, requiring several expensive changes that would also affect performance:

  • Query expressions that drive data changes must be locking
  • Autoincrement must be serialized
  • Implement next-key and end-of-table locks
  • Other stuff

This is not the complete list of changes that we'd have to make to support SBR, nor does it take into account potential enhancements to MySQL replication, but captures the essence of the issue.

(My interest in replication stems from a two-year stint on the replication project at Pervasive Software, however, that was a peer-to-peer implementation employing a change-capture mechanism rather than binlogs--a fundamentally different approach.)

Falcon Schedule

[Disclaimer: I cannot disclose product release dates. Sorry.]

Release planning is a black art, subject to the hazards of miscalculation, fuzzy estimations and plain old bad luck. Our biggest question at the London meeting was how to coordinate the performance optimization effort with that of fixing stability and managing the bug load. Assigning a block of time to any of one of these tasks, much less all three, is not only difficult but outright dangerous if not done carefully. But it had to be done, and so we did.

Assigning priorities is a classic problem in software development, but the textbook approach doesn't always apply. For example, performance optimization should be done late in the project, right? Ok, but what if performance is a major feature, or what if the performance in some cases, is really bad? Isn't that a bug? And what if the performance fix requires a change to a core component, such as the synchronization model or caching algorithm? Wouldn't you want those sort of changes in earlier rather than later?

I suppose with well-defined release criteria and well-defined performance metrics and well-defined defect criteria and well-defined quality metrics, the answers to such questions might be readily apparent, but alas, ours is not the tidy, frictionless universe of textbook theory. No, ours is a messy, non-Newtonian hodgepodge of money, marketing and mission creep, an existential riot fueled by personality, tempered by perception and begging for order in an unordered cosmos.

But I digress.

Actually, we really do have well-defined release criteria and even some pretty good performance criteria, although "good performance" has been something of a moving target for Falcon. The biggest unknown was the time required to bring the bug count under control.

[Disclaimer: I am wholly unqualified to comment upon MySQL 5.1]

After considerable and sometimes heated debate, we did manage to hammer out a shaky consensus on a GA date, saddled with preconditions and "only ifs" as it was. The Falcon schedule has since been revised, and last week in Riga it was thoroughly rehashed again within the context of the Server 6.0 release schedule.

And that's all I have to say about that.

Pilgrimage

The stark reality of the Falcon team meetings is that we spend 80% of the time sitting around a conference table screaming at each other.

No, that's not true. Let me try again.

The reality of our meetings is that we spend a long week of long days in a conference room digging into technical stuff. Despite the travel overhead, I find these meetings to be extremely productive and personally energizing.

In London, there were only two things I absolutely had to do: Visit the British Museum, and, well, see for yourself.

I won't bore you with museum photos, but I did manage to I rope Kevin, Vlad and Vlad's wife, Tatyana, to make the pilgrimage with me to Apple Studios. Tatyana was kind enough to shoot three of us crossing the zebra crossing. Unfortunately, some bearded dude jumped in front of us while we crossed.

He didn't say much. I think he was French.

Saturday, September 27, 2008

Troublemakers, Part I

Philip Stoev, Software Engineer

The mission of the System QA team is to beat the living daylights out of a MySQL release
before it is set free into the world. Philip Stoev is on the System QA team, not the Falcon team, but bless his Bulgarian heart, he's caused more trouble on Falcon than anyone in recent memory except for perhaps Peter Z.

The "trouble" of which I speak is best illustrated by this vignette: Imagine that HMS Falcon is on a shakedown cruise in the Mediterranean. Weather is clear, crew is happy, course is plotted. Midshipman Stoev voices a concern.


"I beg your pardon, Captain, but I believe the ship has run aground."

"Nonsense, Mr. Stoev. The ship is quite sound. She floats upright, the sails are full and the crew is happy with rum and song.
"

"Yes, of course, Captain, but if you'll please have a look belowdecks, sir. The bilges are awash, I'm afraid, and the keel is broken."

"Nonsense, Mr. Stoev, you imagine the worst. Now, hand me my spyglass, if you please, and go on about your duties."

"As you wish, sir."

Get the picture?

Philip does his job quite well such that he can, with the power of a two-line email, stop an entire release dead in its tracks. Last year, this caused tremendous frustration because System QA didn't begin testing until after a release clone-off, which is perilously late in the release cycle. One or two showstopper bugs can, and did, gum up the works for weeks, even months.

Since then we've fixed the process, and now System QA tests are run regularly against pre-release code. There even exists an array of Pushbuild2 servers dedicated solely to System QA stress tests, so each push into the tree results in a battery of stress test executions.

"Executions", indeed. These unassuming, cold-hearted tests, cobbled together from Perl and PHP, efficiently render the very fabric of our precious little engine into hot, screaming shards of digital shrapnel.

Case in point: falcon_online_alter.
This little gem is designed to "exercise" Falcon's online alter capability in the same sense that having your shirt pulled over your head and being kicked down the stairs is "exercise". If I sound bitter, it's only because I've been chasing stress test bugs for the better part of two months.

Let's take an example:

ALTER TABLE t1 ADD INDEX i1 (s1), ADD INDEX i2 (s2);

In MySQL 6.0, "ONLINE" is implicit. Prior to the ONLINE ALTER implementation in Falcon, ADD INDEX worked like this:

1. Get shared lock on table T1

2. Create temporary table #sql-abc with new attribute (e.g. index)
3. Copy records from T1 to #sql-abc (index created during copy)
4. Get exclusive lock on T1
5. Rename T1 to #sql-xyz
6. Rename #sql-abc to T1
7. Drop #sql-xyz

This is a horrendously inefficient way to create an index, so Falcon now creates them online such that a new index is created and populated in place--no temporary tables, no copying, etc. (the same is true with ADD COLUMN, although it's temporarily disabled due to a bug.)


So, with online alter, ADD INDEX works like this:

1. Server queries Falcon to see if it can do the operation online
2. Falcon says, "Yeah, sure, go for it."
3. Server commands online alter
4. Falcon creates new index
5. Falcon populates new index

Falcon briefly locks the table during the DDL phase, but then normal operations resume, even while the index is populated. Pretty cool, right?

Well, sure, except when Philip runs his random query generator that issues thousands of random ALTER operations emanating from many client threads. This is a wholly unnatural act--I mean, what sane application performs thousands of ALTERs from multiple clients? It's nuts, just nuts.

But, and this is a big but, this test finds stuff. All sorts of stuff. Imagine putting a beautiful, Ford 4.6L 3-valve SOHC up on an engine test stand, bolt it down, then run it at 8,000 RPM until it flies apart. Automotive engineers actually do this, so why not software engineers? We need to know where the weak spots are, and, ideally, fix them. And so we do.

Chill/Thaw

Philip also has a chill/thaw stress test, falcon_chill_thaw. This, too, is pure, resident evil, and I mean this in the kindest way.

The Falcon "chill/thaw" mechanism works like this: Each record in the record cache consists of a "system" part and a "data" part. When the total amount of record data associated with an active transaction exceeds a predefined threshold, Falcon "chills" all of the records in the transaction by writing the data portion of each record to the serial log and freeing it from the record cache. When necessary, Falcon "thaws" individual records by restoring their data portions from the serial log.


By default, the falcon_record_chill_threshold is 5MB, which means that when the size of a transaction exceeds 5MB, Falcon moves the record data for the transaction into the serial log. For the falcon_chill_thaw stress tests, Philip cranks the record chill threshold down to 1 byte, which means every record for every transaction is chilled and thawed countless times.

Again, this is a wholly unnatural act, no doubt illegal in the state (not country) of Georgia, and, in fact, I believe we are going to disallow chill thresholds lower than 32K. However, this test has exposed flaws in the chill/thaw code path that ordinarily would never see the light of day.

I have in a logfile a beautiful example of three threads trying to thaw the same record. Falcon is designed to resolve this intrinsic race condition with a compare-and-swap of the data pointer in Record::thaw(). However, the code path leading to the CAS is not (yet) properly reentrant, so we get all sorts of interesting behavior as a consequence of multiple concurrent thaws.

On a live system, this would likely manifest in seemingly random, untraceable errors.
So, it's better to find this stuff now, but it is necessarily time-consuming.

The simple fact remains, however, that Philip is responsible for exposing some serious synchronization issues within Falcon, and for that I am grateful.

I think.

New Falcon Engineers

Apart from a new project manager, the Sun acquisition also gifted our beloved little project with three Sun engineers from DBTG. I will spare the reader personal chagrin and preemptively counter the reflexive incantation of Brooke's Law by stating up front that, in the case of Falcon, you would be wrong. "Mythical Man-Month" my ass, we needed the help.

Naturally, we don't want just anyone hacking Falcon (yes, we're open source, but we have to ship for heaven's sake), and so we didn't get just anyone, we got three top-notch engineers each of whom bring something to the table.

(Free idea: "Heaven's Sake", a tangy Japanese alcoholic beverage made from fermented rice, available at select wine shops near you.)

Olav Sandstaa, Senior Software Engineer

Here is Olav's official bio: "Olav has developed database systems for the last eight years. He worked on the internode communication system, the database kernel as well as performance optimization of HADB. Lately, he has worked as team lead for the performance team for Java DB. He has also made key contributions to design, architecture and technology evaluations for Project Cloudberry. "

"Olav has a master's degree on the topic of parallel execution of relation algebra database operations and a PhD in storage systems for digital video archives from the database research group at the Norwegian University of Science and Technology."

What's this actually means is that Olav--sorry, Doctor Olav--is smarter than you. And me. What this als
o means is that he is really good at Scrabble, Norwegian Database Edition.

The cool thing about Olav is that he doesn't wear his impressive credentials on his sleeve, though he does have "Database God" tattooed behind his left ear (not seen in photo.)

Last week in Riga, I asked Olav what HADB was. He said, "It's like MySQL Cluster, except better." Ouch.


John Embretsen, Software Engineer

Again, I will punt and provide John's official bio:

"John joined Sun in 2005, after graduating from the Norwegian Unive
rsity of Science and Technology. He has been working on Java DB QE/QA, including functional, long-running, usability and compliance testing. He recently gained committer privileges in the Apache Derby community, and have lately been working on adding JMX management and monitoring features to Java DB."

John's official bio fails to mention that he is also the tallest member of the Falcon team (sorry, Kevin), and has an impressive record playing Norwegian basketball, which in the U.S. we call "basketball". At right is a photo of John in a typical just-got-
back-from-rigorous-mountain-hiking-and-now-I-will-sing pose.

Fortunately for us, John accepted a role in Falcon QA, where help was desperately needed.
One of his first tasks was to dig through the suite of testcases relegated to a vile little backwater we call "falcon_team". These are tests that fail for no apparent reason on the pushbuild regression servers, so to keep the matrix green we simply move uncooperative tests out of the way so they can be dealt with later. Well, later is now, because "falcon_team" is our mutant cousin thumping around in the attic, and it's time to walk up those stairs.

Which brings to mind one test that John insists on holding up to the col
d light of day, a test that previously failed intermittently with odd warnings and an error. He now reports that after I implemented the ONLINE ALTER features in Falcon, this test still fails but the row numbers are off by one.

This seemingly innocuous observation carries little meaning to the untrained eye, however, a set of discrete database operations resulting in row numbers being off by one is like finding an extra proton in an atom--you can't simply brush such a thing aside. The magnitude of the discrepancy is small, but the implication of the discrepancy is immense. I am, perhaps, being a bit hyperbolic (hydrochloric?), but the essential fact is that this means more digging into a feature should've been wrapped up weeks ago. Thanks John. Love ya, bro.


Lars-Erik Bjørk, Software Engineer


"Lars-Erik joined Sun in July 2006, shortly after graduating from the Norwegian University of Science and Technology. He has been working 18 months on HADB kernel development and 4 months on PostgreSQL."

Ok, that's pretty cool: HADB kernel development and PostgreSQL experience. Some of you may also recall that Lars-Erik had a previous career in the music industry as the lead vocalist in the group "A-ha", which found brief popularity in the 80's.

Following the dizzying spike of international adoration garnered from the gr
ammatically curious, "Take On Me", Lars-Erik left the band to pursue a career in software engineering. Astute readers will note that Lars-Erik remains physically unchanged from his pop music days, which he attributes to a superior genetic makeup, far-infrared saunas, rock-filtered Norwegian spring water and volcanic mineral baths.

Since joining Falcon in May, Lars-Erik has been working on a variety of bugs. Most recently, (in fact, so recent that he doesn't know it yet) he's working on a couple of chill/thaw bugs and another with the Falcon interface to the Information Schema. (Lars-Erik, if you're reading this, talk t
o Kevin.)



Stalled Thread

Strong start. Zero follow-up. Let's get this going again, shall we?
It's been a busy five months since the last post, so here's an overview of what we've been up to:

May: New Falcon Engineers
Synergy happens. Details in the next post.

May: Falcon Meeting in London
Big news from Jim. Training for the new folks. Details to follow.

June: Falcon 6.0.5 Alpha
Team is busy with bug fixes trying to keep the release train rolling.

July: Falcon Meeting in Boston
Got the band back together, this time on our turf.

August: Falcon 6.0.6 Alpha
Another busy month of bug fixing. Feature complete, finally. Some performance stuff, too.

September: MySQL/Sun Meeting in Riga
All-hands engineering meeting, very productive for Falcon. Details to follow.

Monday, April 28, 2008

The Falcon Team

Ok, one more human interest post before diving into the technical stuff. Let's have a look at the Falcon team as it stands today.

Kevin Lewis, Falcon Team Lead

Kevin joined MySQL two years ago, following a ten-year run as an engineer on the MicroKernel Database Engine team at Pervasive Software, (Btrieve, Pervasive.SQL) in Austin, Texas. He ascended to team leadership when Calvin, our project manager, left MySQL in January.

I worked with Kevin for five years at Pervasive, where he started six months before me, just as he did at MySQL. Kevin's a good friend and has at various times served as something of a father confessor to me, patiently indulging some of the more, shall we say, unsettled phases of my life.

Kevin walks his talk and is supremely ethical. He's a year older than me but looks ten years younger, and twice introduced me to the humbling rigor of backpacking in the Colorado Rockies. (T-shirt idea: "I got b****-slapped by a fourteener.")

Kevin's risen nicely to the challenge of running the Falcon team, and has managed to navigate the turbulent waters of leading a high-visibility team within an open source company. I wouldn't say that leading Falcon is like herding cats--we're not finicky anarchists--but it might be compared to producing the Juilliard senior class play for a performance before the U.N. General Assembly. Or something like that.

Anyway, Kevin's the kind of guy you want to lead a team because he has steady nerves, an assured demeanor, and is physically larger than any of us.


Jim Starkey, Principal Software Engineer, Server Architect

Falcon is Jim's baby, and he is, by all rights, alpha dog on the project. He wrote most of Falcon before MySQL acquired his company, Netfrastructure, in 2006.

I won't provide much of Jim's history here, suffice to say that he's been in the database industry longer than most MySQLers have been alive (seriously), and that he has a proven track record of developing and shipping databases that people pay for and use.

(This is not a minor point. I know many bright, confident and very skilled developers, in and out of the open source world, who lack the practical experience of shipping and supporting a product. If you'll forgive the war metaphor, it's like training for combat but never experiencing it.)

To his credit, and to the occasional dismay of his more reticent counterparts, Jim kicks up a lot of dust inside MySQL, frequently challenging the "It Has Always Been Thus" mindset that tends to accrete in any organization. And, in keeping with the war metaphors, Jim will not hesitate to toss a conversational hand grenade into a quiescent forum or otherwise unremarkable exchange, but only because that's what's on his mind at the time.

All of this is a good thing, in my opinion, because Jim's unsubtle energy is usually countered and productively engaged with corresponding intensity--there is no lack of intensity at MySQL--and some pretty good ideas get tossed about as a result.

Ann Harrison, Lead Software Engineer, Server Architect

Ann inhabits several roles on the Falcon team: architect, developer, documentarian. She is the go-to person for certain elements of the current or future Falcon architecture, Foreign Keys being a good example.

Actually, Foreign Keys is a really good example, because Ann must not only hammer out agreements with the other architectural heavy-hitters at MySQL, she must also comprehend (or convincingly pretend to comprehend) the limitless tracts of arcane Foreign Keys documentation.

Ann is also Cher to Jim's Sonny. If you are unfamiliar with this American cultural reference, let me paint the picture this way. Imagine Jim holding forth on a topic about which he feels strongly, whether on the phone, during a presentation or at the hotel bar. At some point, a clear, authoritative yet uncontentious voice will arise in challenge:

"Uh, no, Jim. Actually, that's not the case at all."

NMI triggered. Handler routine executes.

"Excuse me? What's not the case?"

"You said ABC. It's actually XYZ."

"No. It's ABC. And how on Earth can you possibly know that? I wrote the code myself."

"Yes. And I pointed out to you that it was wrong, and then you fixed it."

"Oh. You're right. Sorry. Ann's right. It's XYZ, not ABC. Anyway, as I was saying..."

Error corrected. Runtime processing resumes.

Trivia: Jim and Ann sail boats. Both are both private pilots. Ann has an instrument rating.


Hakan Küçükyılmaz, Senior Software Engineer, Falcon QA Lead

Hakan has Turkish citizenship, is of Georgian ethnicity and lives in Germany. Like all Europeans, Hakan speaks twenty-two foreign languages. Hakan developed some of his database chops working with (if not for) SAP.

Hakan is responsible for Falcon QA--an enormous challenge--and has been at MySQL longer than any of the Falcon team, since 1948, I think.

Apart from adapting the existing InnoDB test cases, the real challenge for Hakan has been to develop test cases that exercise elements of the Falcon engine that are uniquely Falcon--high concurrency and a heavy reliance on abundant RAM, for instance.

Hakan's efforts have been supremely valuable in helping us pin down exactly what this Falcon thing is relative to other storage engines and relative to what we expect it to be. Last summer, for example, Hakan started a weekly DBT2 run which, after no small amount of haggling over the test configuration, eventually became the primary engineering measure of Falcon performance.

[N.B. I must emphasize here that we use DBT2 as an internal metric with which to measure Falcon performance relative to InnoDB. It is by no means regarded as the best representation of Falcon performance, and, in fact, was chosen precisely because it tends to favor InnoDB, especially at low concurrencies. However, DBT2 is a measure of something, and judging by the number and quality of bugs resolved as a consequence of its use, DBT2 has proven to be an outstanding tool for evaluating Falcon performance.]

With the recent Sun acquisition, we now have the kind attention of the Sun Performance and Availability team, whose sole charter is to measure and improve performance on Sun hardware, which brings to mind a quote from Terminator: "That's what he does. That's all he does! You can't stop him!"

Anyway, I like Hakan. He is outspoken. He keeps us real. He's probably the hardest working member of the team. Hakan brings to the Falcon team a kind of weary, fatalistic edge that serves as a fitting balance to my own eternal optimism.

Imagine Friedrich Nietzsche. Now imagine Eeyore. There you have it.

The picture is of Hakan doing something really cool, I'm just not sure what.


Vladislav Vaintroub, Senior Software Engineer

Vlad joined Falcon in January. He's Russian, lives in Germany, has a math degree. Like most Europeans, Vlad speaks twenty-two foreign languages, many of them English.

Vlad got much of his database experience at Software AG. At MySQL, he hit the ground running and was fixing Falcon bugs before he got his first paycheck.

Vlad bravely or naively took over the Supernodes feature, a kind of meta-index (Kevin doesn't like that description, but I'm sticking to it.) He reworked the Supernodes code a bit, fixed the bugs, checked it in, demonstrated it to our beloved VP of Engineering and finally measured a significant performance improvement (more on performance later.) Not bad for his first three months on the job.

Vlad's soft-spoken and unobtrusive nature has a somewhat calming influence on the team, I think. He also possesses a rare talent (rare in the context of MySQL) that is shared by all Falcon developers, which is the ability to develop on Windows.

At right is Vlad affecting ironic detachment. The photo was taken shortly after the other Beatles had crossed the road.

Kelly Long, Senior Software Engineer, Performance Architect

I said this to Kevin at the MySQL Users Conference:

"I had a great discussion about Falcon performance with Kelly last night at the bar. This guy's a sleeper, man, I'm telling you."

First, just let me say that all of my great conversations at the Users Conference took place at the bar. Not so much because of the drinking (no, really), but because that's where everyone lands at the end of the day.

Kelly Long works for Mikael Ronström's architecture group and is not, strictly speaking, a member of the Falcon team. However, since late last year, Kelly's primary focus has been to measure and analyze Falcon performance. To give you a sense of how seriously he takes this job, at right is a snapshot of the gear this cat has at his home.

My remarks to Kevin that evening were, in part, inspired by some outstanding analysis Kelly did on the Falcon SyncObject implementation, and the subsequent fix. I'll save the details for a near-future post, but the fix closed a timing window whereby sleeping threads were not awakened when the object for which they were waiting became available.

The problem typically manifested at higher concurrencies (64 to 256 clients) on an 8-way system. The symptom was a perplexing and, frankly, frightening standard deviation. For example, at 64 clients, InnoDB might exhibit an STD of 1.8% whereas Falcon would sometimes have an 80% STD, rendering the results effectively worthless.

Apart from his recent SyncObject discovery, Kelly had also performed in-depth analysis on other sacrosanct components of the Falcon core, including some of the page cache algorithms, and was in the process of trying out code changes to verify his conclusions.

Now, in the context of our small and busy little team, I regarded this news as an absolute gift, because it was precisely the type of analysis we've needed but had neither the time nor the resources to perform.

I should also note that, as a seasoned senior engineer, Kelly had the good sense to establish some empirical evidence before announcing his findings, because changes to the Falcon core are not considered lightly, and we just don't have time to chase every "what-if" without probable cause unless it exhibits prima facie evidence of potential awesomeness.

Manyi Lu, Engineering Manager, MySQL Server 6.0

Manyi Lu is an engineering manager from the Sun Database Technology Group. She is from Beijing, lives in Norway and speaks at least two more languages than me. Manyi very recently became project manager of the MySQL Server 6.0 and related groups, including Falcon, Maria, Backup and the Subquery optimization teams.

Manyi introduced herself to Kevin and me at Marten's party during the UC. At the time, I was blathering on to Kevin and Kelly about some damn thing or another, when Manyi and Rune Humborstad, a Sun software architect, drifted over to us and started asking questions about Falcon's stability and performance.

I did not know Manyi or Rune, but was intrigued by their interest in Falcon's stability, which was a bit rocky at the time. Kevin and I were quite happy to oblige, and spoke candidly about the project, neither of us realizing that Manyi would soon become the MySQL Server 6.0 project manager.

We were sort of awkardly positioned around a pesky little five-foot tree in Marten's back yard, but it was a very interesting and positive conversation, and Manyi offered to have some of the DBTG group help with Falcon QA and possibly bug fixing. To me, this was the first tangible evidence of how the Sun acquistion can make a real difference in the outcome and success of the Falcon project.

I left Marten's party really psyched.



The MySQL Falcon team at Heidelberg Castle, September 2007 .

L-R: Hakan, Kevin, Calvin, Ann, Jim, Christoffer Hall (now on the Cluster team), Chris Powers.