[pca] Cluster, zones, noreboot.

David Stark dave at davidstark.name
Tue Jul 13 18:35:23 CEST 2010


Hi Enda, Don, Chris.

On 13/07/2010 16:35, Enda O'Connor wrote:
> Hi
>
> Some people have recommended Update On Attach see
> http://wikis.sun.com/display/BluePrints/Maintaining+Solaris+with+Live+Upgrade+and+Update+On+Attach
>
>
> and also the following for a description of how update on attach works
> in conjunction with patching.
> http://www.sun.com/bigadmin/features/articles
> /zone_attach_patch.jsp#Patching

For this round of patching, we'll have to bring everything down to 
patch, since we don't have 137137-09 installed so we can't use Update On 
Attach yet. The next round of patching could be different, though.

Having read that BigAdmin article, though, I'm a bit wary - increasing 
NGZs' root filesystem usage isn't very desirable in our situation and 
the prospect of having some packages potentially left un-patched doesn't 
sound too wonderful either (auditors abound here...).

> it is important to first read the bigadmin article to understand how it
> works before goign down this route.
>
> Enda
> On 13/07/2010 16:32, Enda O'Connor wrote:
>> Hi David
>> the major issue with patching a live system with a failover zone is if
>> the zone failed over for any reason during patching.

OK. I think I would probably remove all the other hardware nodes from 
each zone's resource group's nodelist before trying this, to prevent the 
zones from moving.

>>This would cause
>> patch corruption, one would need to suspend the HA container resource,
>> i.e.
>>
>> clrg suspend << the resource group >>
>> detach zones on remaining node
>> apply patch
>> attach zone on other node
>> clrg resume the resource group
>>
>> But one would need to take some care to identify patches that can be
>> applied in such fashion.

I'm hoping the --noreboot flag in PCA should cover this. Patches without 
any of 'reconfigimmediate', 'rebootimmediate', 'reconfigafter' or 
'rebootafter' in their PATCH_PROPERTIES are classed as noreboot patches 
by PCA - should that be enough in a Cluster environment?

>> the following doc has section on applying patches that require Single
>> User Mode in failover zone environment.
>> http://docs.sun.com/app/docs/doc/819-2971/z4000076997776?a=view
>>
>> I have cc'ed Chris who has lots of experience in this area.

Ta!

>> But the main concern is that the zone might failover during such
>> patching.
>>
>> Enda

Cheers!

Dave

>> On 13/07/2010 10:56, Don O'Malley wrote:
>>> Hey Enda/Ed,
>>>
>>> Any thoughts on this?
>>>
>>> In addition to the no reboot question, is it better to detaches your
>>> zones and use update on attach to bring the local zones back in sync, or
>>> is there no difference between the two (I thought update on attach was
>>> quicker)?
>>>
>>> Best,
>>> -Don
>>>
>>>
>>> David Stark wrote:
>>>> Hi List.
>>>>
>>>> A bit off-topic, but PCA's involved, so I'm going to push my luck.
>>>>
>>>> We've got a number of Sun Cluster installs using (lots of) failover
>>>> zones and with no LiveUpgrade alt. boot space set up, which we need to
>>>> patch. If we were to patch the clusters node-by-node (including the
>>>> kernel patches that don't like zones being booted when they're
>>>> applied), the zones would fail over between the nodes and never get
>>>> patched themselves, and since kernel (and some other?) patches can't
>>>> be applied from inside zones we would end up in a situation where the
>>>> zones' patch databases are out of sync with the Global zone. I know
>>>> from very bitter experience that this is a Bad Thing, so to avoid that
>>>> we're currently bringing the clusters down to patch them. This
>>>> obviously isn't optimal - people expecting 100% uptime from the
>>>> clusters are naturally a bit annoyed at having their applications down
>>>> for several hours while we unleash the mighty PCA.
>>>>
>>>> So, to minimise downtime I'd like to apply the noreboot patches, say,
>>>> the night before, and have a more minimal patch run with the clusters
>>>> down. This brings me to the question:
>>>>
>>>> Has anyone ever had any problems with noreboot patches applied to live
>>>> systems? Any weirdness at all? I've patched plenty of test machines in
>>>> multi-user mode, but never busy production boxes - these are fairly
>>>> large Oracle and SAP environments for the most part.
>>>>
>>>> Anyone with experience patching Sun Cluster care to share any top tips?
>>>>
>>>> Cheers!
>>>>
>>>> Dave
>>>>
>>>
>>> --
>>> <http://www.oracle.com/>
>>> *Don O'Malley*
>>> Manager,Patch System Test
>>> Revenue Product Engineering | Solaris | Hardware
>>> East Point Business Park, Dublin 3, Ireland
>>> Phone: +353 1 8199764
>>> Team Alias: rpe_patch_system_test_ww at oracle.com
>>> <mailto:rpe_patch_system_test_ww at oracle.com>
>>> <http://www.oracle.com/commitment>
>>
>




More information about the pca mailing list