NotificationsYou must be signed in to change notification settings
Fork6
Star31

Commitf155466

committed

Fix bogus concurrent use of _hash_getnewbuf() in bucket split code.

_hash_splitbucket() obtained the base page of the new bucket by calling_hash_getnewbuf(), but it held no exclusive lock that would prevent someother process from calling _hash_getnewbuf() at the same time. This iscontrary to _hash_getnewbuf()'s API spec and could in fact cause failures.In practice, we must only call that function while holding write lock onthe hash index's metapage.An additional problem was that we'd already modified the metapage's bucketmapping data, meaning that failure to extend the index would leave us witha corrupt index.Fix both issues by moving the _hash_getnewbuf() call to just before wemodify the metapage in _hash_expandtable().Unfortunately there's still a large problem here, which is that we couldalso incur ENOSPC while trying to get an overflow page for the new bucket.That would leave the index corrupt in a more subtle way, namely that someindex tuples that should be in the new bucket might still be in the oldone. Fixing that seems substantially more difficult; even preallocating asmany pages as we could possibly need wouldn't entirely guarantee that thebucket split would complete successfully. So for today let's just dealwith the base case.Per report from Antonin Houska. Back-patch to all active branches.

1 parentd12afe1 commitf155466Copy full SHA for f155466

File tree

1 file changed

+26

-4

lines changed

src/backend/access/hash
- hashpage.c

1 file changed

+26

-4

lines changed

`‎src/backend/access/hash/hashpage.c‎`

Lines changed: 26 additions & 4 deletions

Original file line number	Diff line number	Diff line change
`@@ -37,6 +37,7 @@`
`37`	`37`	`staticbool_hash_alloc_buckets(Relationrel,BlockNumberfirstblock,`
`38`	`38`	`uint32nblocks);`
`39`	`39`	`staticvoid_hash_splitbucket(Relationrel,Buffermetabuf,`
	`40`	`+Buffernbuf,`
`40`	`41`	`Bucketobucket,Bucketnbucket,`
`41`	`42`	`BlockNumberstart_oblkno,`
`42`	`43`	`BlockNumberstart_nblkno,`
`@@ -176,7 +177,9 @@ _hash_getinitbuf(Relation rel, BlockNumber blkno)`
`176`	`177`	`*EOF but before updating the metapage to reflect the added page.)`
`177`	`178`	`*`
`178`	`179`	`*It is caller's responsibility to ensure that only one process can`
`179`		`- *extend the index at a time.`
	`180`	`+ *extend the index at a time. In practice, this function is called`
	`181`	`+ *only while holding write lock on the metapage, because adding a page`
	`182`	`+ *is always associated with an update of metapage data.`
`180`	`183`	`*/`
`181`	`184`	`Buffer`
`182`	`185`	`_hash_getnewbuf(Relationrel,BlockNumberblkno,ForkNumberforkNum)`
`@@ -503,6 +506,7 @@ _hash_expandtable(Relation rel, Buffer metabuf)`
`503`	`506`	`uint32spare_ndx;`
`504`	`507`	`BlockNumberstart_oblkno;`
`505`	`508`	`BlockNumberstart_nblkno;`
	`509`	`+Bufferbuf_nblkno;`
`506`	`510`	`uint32maxbucket;`
`507`	`511`	`uint32highmask;`
`508`	`512`	`uint32lowmask;`
`@@ -615,6 +619,13 @@ _hash_expandtable(Relation rel, Buffer metabuf)`
`615`	`619`	`}`
`616`	`620`	`}`
`617`	`621`
	`622`	`+/*`
	`623`	`+ * Physically allocate the new bucket's primary page. We want to do this`
	`624`	`+ * before changing the metapage's mapping info, in case we can't get the`
	`625`	`+ * disk space.`
	`626`	`+ */`
	`627`	`+buf_nblkno=_hash_getnewbuf(rel,start_nblkno,MAIN_FORKNUM);`
	`628`	`+`
`618`	`629`	`/*`
`619`	`630`	`* Okay to proceed with split. Update the metapage bucket mapping info.`
`620`	`631`	`*`
`@@ -668,7 +679,8 @@ _hash_expandtable(Relation rel, Buffer metabuf)`
`668`	`679`	`_hash_droplock(rel,0,HASH_EXCLUSIVE);`
`669`	`680`
`670`	`681`	`/* Relocate records to the new bucket */`
`671`		`-_hash_splitbucket(rel,metabuf,old_bucket,new_bucket,`
	`682`	`+_hash_splitbucket(rel,metabuf,buf_nblkno,`
	`683`	`+old_bucket,new_bucket,`
`672`	`684`	`start_oblkno,start_nblkno,`
`673`	`685`	`maxbucket,highmask,lowmask);`
`674`	`686`
`@@ -751,10 +763,16 @@ _hash_alloc_buckets(Relation rel, BlockNumber firstblock, uint32 nblocks)`
`751`	`763`	`* The caller must hold a pin, but no lock, on the metapage buffer.`
`752`	`764`	`* The buffer is returned in the same state. (The metapage is only`
`753`	`765`	`* touched if it becomes necessary to add or remove overflow pages.)`
	`766`	`+ *`
	`767`	`+ * In addition, the caller must have created the new bucket's base page,`
	`768`	`+ * which is passed in buffer nbuf, pinned and write-locked. The lock`
	`769`	`+ * and pin are released here. (The API is set up this way because we must`
	`770`	`+ * do _hash_getnewbuf() before releasing the metapage write lock.)`
`754`	`771`	`*/`
`755`	`772`	`staticvoid`
`756`	`773`	`_hash_splitbucket(Relationrel,`
`757`	`774`	`Buffermetabuf,`
	`775`	`+Buffernbuf,`
`758`	`776`	`Bucketobucket,`
`759`	`777`	`Bucketnbucket,`
`760`	`778`	`BlockNumberstart_oblkno,`
`@@ -766,7 +784,6 @@ _hash_splitbucket(Relation rel,`
`766`	`784`	`BlockNumberoblkno;`
`767`	`785`	`BlockNumbernblkno;`
`768`	`786`	`Bufferobuf;`
`769`		`-Buffernbuf;`
`770`	`787`	`Pageopage;`
`771`	`788`	`Pagenpage;`
`772`	`789`	`HashPageOpaqueoopaque;`
`@@ -783,7 +800,7 @@ _hash_splitbucket(Relation rel,`
`783`	`800`	`oopaque= (HashPageOpaque)PageGetSpecialPointer(opage);`
`784`	`801`
`785`	`802`	`nblkno=start_nblkno;`
`786`		`-nbuf=_hash_getnewbuf(rel,nblkno,MAIN_FORKNUM);`
	`803`	`+Assert(nblkno==BufferGetBlockNumber(nbuf));`
`787`	`804`	`npage=BufferGetPage(nbuf);`
`788`	`805`
`789`	`806`	`/* initialize the new bucket's primary page */`
`@@ -832,6 +849,11 @@ _hash_splitbucket(Relation rel,`
`832`	`849`	`* insert the tuple into the new bucket. if it doesn't fit on`
`833`	`850`	`* the current page in the new bucket, we must allocate a new`
`834`	`851`	`* overflow page and place the tuple on that page instead.`
	`852`	`+ *`
	`853`	`+ * XXX we have a problem here if we fail to get space for a`
	`854`	`+ * new overflow page: we'll error out leaving the bucket split`
	`855`	`+ * only partially complete, meaning the index is corrupt,`
	`856`	`+ * since searches may fail to find entries they should find.`
`835`	`857`	`*/`
`836`	`858`	`itemsz=IndexTupleDSize(*itup);`
`837`	`859`	`itemsz=MAXALIGN(itemsz);`

0 commit comments

Comments

(0)

Movatterモバイル変換

Navigation Menu

Search code, repositories, users, issues, pull requests...

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Commitf155466

File tree

1 file changed

1 file changed

`‎src/backend/access/hash/hashpage.c‎`

0 commit comments